SparseDFF: Sparse-View Feature Distillation for One-Shot Dexterous Manipulation
Qianxu Wang, Haotong Zhang, Congyue Deng, Yang You, Hao Dong, Yixin Zhu, Leonidas Guibas
Introduction
Learning from demonstration is an effective means of rapidly endowing robotics with complex skills. Yet to date, to proficiently transfer manipulation skills to robots, a series of meticulously designed demonstrations, enriched with diverse scene and task variations and bolstered by data augmentation, are often essential. Conversely, humans exhibit extraordinary extrapolation and generalization capabilities when learning from demonstrations. For example, acquiring the skill to hold a cat by observing another person can likely be generalized to holding various other cats with different breeds, shapes, and appearances, or even to different animals like dogs, otters, or baby tigers, assuming they are approachable and cooperative. The crux of such robust generalization lies in understanding the inherent similarities of these instances despite the variations in appearances, poses, or categories.
To equip autonomous agents with comparable high-level understandings, an effective approach is leveraging object and scene representations from large vision models. Given that most of these models are trained on 2D images—due to the challenges inherent in 3D data acquisition and annotation—it poses a significant challenge to directly apply them for complex manipulation tasks, such as those involving a dexterous hand. Recent research has introduced DFFs (Kobayashi et al., 2022), which convert dense 3D fields from 2D image features, thereby improving 3D scene understanding (Kerr et al., 2023) and interactions (Shen et al., 2023; Lin et al., 2023; Rashid et al., 2023; Ze et al., 2023).
However, most existing studies on 2D-to-3D feature distillation primarily rely on dense camera views. This reliance limits many interaction scenarios where only sparse camera views are available, without the possibility of repositioning the cameras. In such cases, directly extrapolating the DFF similarly to dense-view setups—using feature averaging or interpolation—results in a degradation of feature quality. This degradation stems from the absence of enforcement of view consistency on image features, causing inconsistencies in the overlapping regions between different views.
In this paper, we present SparseDFF, a 3D DFF model learned from sparsely sampled RGBD observations, with as few as 4 views, enabling the one-shot transfer of dexterous manipulations. Given a 3D scene and sparsely sampled camera scans, we initially extract the DINO features (Caron et al., 2021) from each image individually and back-project them to the partial point cloud corresponding to each view. However, due to the sparse nature of the views and the inherent lack of view coherency in DINO features, direct averaging of features from multiviews would lead to a degradation of 3D feature quality. To mitigate this, we introduce a lightweight feature refinement network, trained with a contrastive loss on pairwise views. This network is tailored to a single scene but is designed to be directly applicable to novel scenes without requiring modifications or fine-tuning.
After integrating the partial point clouds from all views with refined features, we employ a point pruning mechanism to further enhance local feature continuity. Subsequently, we propagate the point features throughout the 3D space to formulate a continuous field. The resulting feature fields establish dense correspondences between varying scenes, allowing the definition of an energy function on the end-effector pose between the source demonstration and target manipulation. This enables the learning of dexterous manipulations from one-shot demonstrations, adaptable to novel scenes with variations in object poses, deformations, scene contexts, or even different object categories. We validate our method by real-world experiments involving a dexterous hand interacting with both rigid and deformable objects, showcasing robust generalization across diverse objects and scene contexts.
To summarize, our contributions are threefold: (i) We propose a novel framework for one-shot learning of dexterous manipulations, utilizing semantic scene understandings distilled into 3D feature fields. (ii) We develop an efficient methodology for extracting view-consistent 3D features from 2D image models, incorporating a lightweight feature refinement network and a point pruning mechanism. This enables the direct application of our network to novel scenes to predict consistent features without any modifications or fine-tuning. (iii) Extensive real-world experiments with a dexterous hand validate our method, demonstrating robustness and superior generalization capabilities in diverse scenarios.
Related Work
Identifying point-wise correspondences can aid the transfer of manipulation policies between diverse objects. Contrary to key-point-based methods (Manuelli et al., 2019; Xue et al., 2023; Florence et al., 2018), recent work has concentrated on constructing dense feature fields using implicit representations (Simeonov et al., 2022; Shafiullah et al., 2022; Ryu et al., 2022; Dai et al., 2023; Simeonov et al., 2023; Weng et al., 2023; Urain et al., 2023). With the advancement of large vision models, emerging research is leveraging feature fields distilled from pre-trained vision models to facilitate few-shot or even one-shot policy learning or 6-DoF grasps (Shen et al., 2023), sequential actions (Lin et al., 2023), or language-guided manipulations (Rashid et al., 2023). However, these studies either inherit the limitations of existing 3D feature field distillation methods, requiring dense view inputs (Shen et al., 2023; Rashid et al., 2023), necessitating moving the camera around the entire scene, or rely on features from a single view (Lin et al., 2023), suitable for parallel grippers but less so for dexterous manipulations with increased spatial complexities.
Particularly relevant is Ze et al. (2023), which utilizes several camera views but synthesizes unseen views using neural rendering and extracts features from pre-trained StableDiffusion. This represents progress in bridging sparse and dense observations, but synthesizing and propagating unseen views remain effort-intensive. Additionally, these studies primarily address simple manipulations with parallel grippers, leaving complex end-effector dexterous manipulations unexplored.
Another pertinent work is Karunratanakul et al. (2020), utilizing a signed distance field to depict interactions between human hands and objects. We employ a similar approach as Karunratanakul et al. (2020) to optimize hand parameters from the 3D field; however, unlike their direct representation of hand-object distances, we construct an implicit feature field and use it to define energy functions for the end-effector parameters.
Distilling 2D features into 3D
Zhi et al. (2021), Siddiqui et al. (2023), and Ren et al. (2022) lift semantic information from segmentation networks into 3D, illustrating that averaging potentially noisy language embeddings over input views can yield clear 3D segmentations. Kobayashi et al. (2022) and Tschernezki et al. (2022) investigate embedding pixel-aligned image feature vectors from LSeg or DINO (Caron et al., 2021) into 3D Neural Radiance Fields (NeRF), demonstrating their utility in 3D manipulations of underlying geometry. Peng et al. (2023) and Kerr et al. (2023) further showcase the distillation of non-pixel-aligned image features (e.g., CLIP (Radford et al., 2021)) into 3D without fine-tuning. However, these works utilize dense camera views to derive 3D features. While suitable for 3D understanding, this approach is incompatible with many interaction scenarios where only a few cameras are available at fixed locations.
Dexterous grasping
Dexterous hands have been a focal point of study in the robotics community due to their ability to execute complex actions and their potential for human-like grasping (Salisbury & Craig, 1982; Rus, 1999; Okamura et al., 2000; Dogar & Srinivasa, 2010; Andrews & Kry, 2013; Dafle et al., 2014; Kumar et al., 2016b; a; Nagabandi et al., 2020; Lu et al., 2020; Qi et al., 2023; Liu et al., 2022). Within this domain, dexterous grasping has emerged as a critical area of focus owing to its pivotal role in most hand-object interactions.
Classical analytical methods (Bai & Liu, 2014; Dogar & Srinivasa, 2010; Andrews & Kry, 2013) model the kinematics and dynamics of both hands and objects directly, incorporating varying levels of simplifications. In contrast, recent advancements have been observed in learning-based methods. These include state-based learning, assuming that the robot has access to all oracle states (Chen et al., 2022; Christen et al., 2022; Rajeswaran et al., 2017; Nagabandi et al., 2020; Huang et al., 2021; Andrychowicz et al., 2020; She et al., 2022; Wu et al., 2023), and vision-based methods (Mandikal & Grauman, 2021; 2022; Li et al., 2023; Qin et al., 2023; Xu et al., 2023; Wan et al., 2023) which consider more realistic scene observations.
However, a common limitation of these works is their reliance on extensive demonstration data for training, and many exhibit constrained generalization capabilities. A work closely related to ours is Wei et al. (2023), which introduces a functional grasp synthesis algorithm utilizing minimal demonstrations. Relying on geometric correspondences and physical constraints, Wei et al. (2023) achieves generalizations primarily within object categories featuring highly similar shapes.
In contrast, our method extends the capabilities by leveraging semantic visual features, enabling cross-category generalizations with a single demonstration, providing a more versatile and generalized approach to dexterous manipulation and grasping.
Method
This formulation results in a continuous feature field in the 3D space surrounding the scene.
Addressing the issue of local feature discrepancies, we devise a lightweight feature refinement network , consisting of a two-layer per-point MLP, as depicted on the left in Fig. 2. This network can be efficiently self-supervisedly trained on a single source scene, and we show its efficacy in delivering high-quality consistent features when directly applied to novel scenes without modifications. The effectiveness of this feature refinement is shown in Sec. 4.3.
Given any pair of partial point clouds from different views with per-point DINO features , the application of yields the refined features . The intuition is to ensure that features within the same neighborhood are similar, while those from distant neighborhoods are distinct, aligning with the principles in Xie et al. (2020).
To achieve this, we employ the contrastive learning framework from Chen et al. (2020) to optimize the weights in . The refined features are passed through a projection head to obtain projected features . The contrastive learning objective is defined by identifying the overlapping region of the two views with pairs of points distanced less than 1cm. During each training iteration, a minibatch of pairs is randomly sampled as positive examples, with the other possible pairs of unmatched points within the minibatch serving as negative examples. The contrastive loss for a positive pair of examples is expressed as
where represents the cosine similarity, is an indicator function evaluating to 1 iff , and is a temperature hyperparameter. Post-training on the source scene, the projection head is discarded, and only is applied to novel scenes.
2 Point Pruning
To enhance feature consistency within the merged point cloud , we implement a pruning mechanism. This mechanism operates based on the refined point features and is driven by the feature similarity among proximate points. More precisely, we devise a voting mechanism, illustrated on the right in Fig. 2.
and we discard the 20% of points that accumulate the fewest votes.
This pruning mechanism is crucial as DINO features within the same image are typically locally continuous, but discontinuities may arise when two adjacent points originate from different views. By employing this pruning approach, we retain the point features that garner more consensus from varying views, thereby promoting feature consistency and reliability. This process also takes into account the point cloud noise and the imperfection of DINO features.
3 End-Effector Optimization
Our end-effector, a dexterous hand, is represented by a set of parameters denoting all joint poses. Given a source demonstration consisting of a scene point cloud and hand parameters , and a target scene point cloud , we aim to optimize the hand parameters using the feature fields of both scenes; see also Fig. 3.
To maintain the physical feasibility of the learned action, we incorporate repulsion energy functions from Wang et al. (2023) to avoid hand-object and hand self-penetrations. The energy functions for inter-penetration and self-penetration are defined as:
Additionally, to prevent extreme hand poses that could potentially damage the robot, we impose a pose constraint penalizing out-of-limit hand pose. The overall loss function combines all the aforementioned terms:
In our implementation, we set .
Experiments
We evaluate our model by deploying it to a robot hand and conducting experiments under various setups. Given the superior stability of large image models like DINO (Caron et al., 2021; Oquab et al., 2023) on real photos compared to synthetic images, we opt to assess our method directly in real-world settings, bypassing simulations.
Our experiments utilize a Shadow Dexterous Hand, featuring a total of 24 Degrees of Freedom (DoF). We lock the 2 DoFs associated with the wrist, allocating the remaining 22 for optimization. This hand is attached to a UR10e arm, which has 6 DoFs. The robot operates on tables of dimensions 1m1.2m or 1m1m, with the robotic arm strategically positioned to access any location on the tabletop. However, to ensure equipment safety, the dexterous hand is restricted from making direct contact with the table surface, and the hand’s elbow also imposes limitations on its range of motion.
For acquiring RGBD scans, we employ four Azure Kinect DK sensors, each stationed at a corner of the table, all focused toward the table’s center. These sensors, pre-calibrated and set at fixed elevations, facilitate the capture of the environment. The acquired scans undergo post-processing to eliminate background interferences. In experiments involving a single object, we leverage SAM (Kirillov et al., 2023) to segment the object’s point cloud. For multi-object scene experiments, we employ physical information to exclude areas beyond the tabletop experimental region and implement RANSAC (Fischler & Bolles, 1981) to filter out the table surface.
Tasks and evaluation
We conduct a quantitative evaluation of our method focusing on the grasping of both rigid and deformable objects, as grasping presents clear and definable criteria for success and failure. In each trial, an initial demonstration is provided on an object scan within a virtual environment. This is achieved by manually positioning a dexterous hand on the object point cloud using MeshLab. Subsequently, our feature network is overfit to this source scene, a process that approximately takes 300 seconds. Post this training phase, the network is frozen and deployed to various real-world scenes to perform hand pose optimization, requiring roughly 20 seconds per scene.
For every test scene, object poses are randomized within a reachable range by the hand, and our method is executed 10 times to ascertain the success rate. Each test run initiates with the acquisition of an initial grasping pose through our end-effector optimization, followed by the execution of a predefined lifting motion—either upright or a combination of upward and retreating movements, contingent on the physical constraints. A run is deemed successful if the object is lifted off the table without falling. Beyond these quantitative evaluations, we also showcase qualitative results, extending beyond grasping, to illustrate our method’s proficiency in replicating the action styles observed in the demonstrations.
Baselines
Our main baseline is a naive DFF, wherein the image DINO features are directly back-projected to the point cloud, as illustrated in Peng et al. (2023), and subsequently interpolated throughout the 3D space. Given the nascent state of research in one-shot learning of dexterous manipulation, this baseline employs the same end-effector optimization as our method. For experiments involving rigid objects, we also draw comparisons with UniDexGrasp++ (Wan et al., 2023), a leading method in learning dexterous grasping. Given the instability of UniDexGrasp++’s vision-based model with our noisy point cloud scans, we assess its state-based model within simulation environments on objects possessing digital twins, such as Box1, Drill, Bowls, and Mugs. For each task setup, we conduct 100 tests (as opposed to 10) to determine the success rate.
1 Rigid Objects
We first evaluate our method on rigid object grasping to illustrate the adaptability of our method to varying poses, shapes, and object categories. The evaluation is structured around the following specific setups:
Box: A demonstration is provided, showcasing the grasping of the Cheez-It box from the YCB dataset (Calli et al., 2015b; 2017; a) (ID=3). The evaluation is extended to the same box and another cracker box, each exhibiting distinct geometry and appearance.
Drill: A functional grasp of the Drill from the YCB dataset (ID=35) by its handle serves as the demonstration. The evaluation involves testing on the same drill presented in varied poses.
Bowl: Utilizing a 3D printer, we created three distinct bowls. A grasp by the rim on Bowl1 is demonstrated, followed by testing on all three bowls, each in varied poses. This includes a cat-shaped bowl, notable for its non-standard shape and intricate geometry.
Bowl Mug: The learned grasping from Bowl1 is further transferred to three differently 3D-printed Mugs. This illustrates the method’s capability to generalize to a different object category—Mugs, which, while semantically and functionally related to Bowls, possess disparate geometries.
Tab. 1 presents the success rates of our method compared to the baselines for rigid object grasping tasks. Our approach outperforms across all setups. The baseline methods exhibit commendable results in simpler scenarios, particularly where the manipulation target aligns precisely with the demonstrated object. However, their efficacy diminishes significantly when the source and target objects diverge, notably in challenging transitions like Bowl1 to CatBowl, or Bowl1 to Mugs. In contrast, our method maintains high success rates in these instances.
Fig. 4 presents our qualitative results. On the left, we display the Boxes and Drill from the YCB dataset, illustrating our method’s generalization capabilities under rigid transformations of the object poses. We also depict a transfer from Box1 to Box2, highlighting adaptability to changes in geometry and appearance. On the right, we demonstrate learned grasping from a demonstration on Bowl1, showcasing its generalization to different bowls, including a challenging CatBowl with intricate geometry. Additionally, we exhibit our method’s cross-category generalization capabilities from Bowls to Mugs. Further qualitative results are available in the Supplementary Video.
2 Deformable Objects
We assess our method’s efficacy on deformable object grasping to demonstrate its generalization capabilities across deformations, varying objects, and alterations in scene contexts. The evaluation setups are detailed below:
SmallBear, BigBear, Monkey: Three deformable toy animals are utilized for evaluation, encompassing two differently sized, less deformable Bears and a more flexible, skinny Monkey with elongated limbs. Demonstrations are provided on a fixed pose of each toy, and evaluations are conducted on varying poses and deformations.
Monkey SmallBear: The method’s ability to transfer grasping knowledge between the Monkey and the SmallBear is also assessed, representing a more challenging task requiring both inter-object and deformation generalizations.
Monkey in Context: A scene setup is designed incorporating the Monkey and several background objects, with object positions randomized in each execution. This setup includes challenging cases where the Monkey is occluded by background objects. The demonstration grasp is performed on a single Monkey in this setup.
The success rates are presented in Tab. 2. Our method significantly surpasses the baseline, achieving exemplary success rates, particularly in more challenging scenarios necessitating extensive generalizations. This performance aligns with the results obtained in rigid object evaluations.
Fig. 5 illustrates our qualitative results. In the top left of the figure, the grasp of a SmallBear toy by its body is depicted, demonstrating the method’s generalization to different poses and to a Monkey toy, showcasing the versatility of our approach. The bottom left presents an example of grasping a BigBear by its nose. This specific action is not attempted on other toys as their noses are too small for effective feature extraction and grasp execution, illustrating the method’s adaptability to the physical constraints of different objects.
On the right, the figure displays grasps on the Monkey toy, highlighting the method’s adaptability to more extensive deformations on its limbs. It also showcases a transfer from the Monkey to the SmallBear, emphasizing the method’s capability to generalize across different object forms and structures. At the bottom right, a more challenging test setup is presented with the Monkey placed amidst various scene contexts. A successful grasp in this scenario necessitates not only the identification of the target object but also an understanding of its pose, deformation, and inter-object relations, even in challenging cases of occlusions.
These examples underscore the method’s robustness and adaptability in diverse and challenging scenarios, demonstrating its potential applicability in real-world, dynamic environments. Additional qualitative results and discussions are available in the Supplementary Video.
Our framework extends beyond grasping, demonstrating versatility in various hand-object interactions. Fig. 6 illustrates two distinct interaction scenarios with toy animals, highlighting the adaptability of our method. On the left, the interaction involves caressing the head of a Monkey toy, showcased in different body poses. In the demonstration, the Monkey is lying on the table with its arms stretched straight, while in the test scenes, it is depicted hugging a BigBear with its arms, adapting to the new interaction scenario seamlessly. On the right, the interaction is focused on patting the toys on their butts. This action is demonstrated on the Monkey in the source scene, and in the target scenes, it is generalized to the SmallBear, whether it is alone or placed in a context surrounded by other objects, emphasizing the method’s capability to adapt to varying scenarios.
3 Ablation Studies
In Fig. 7, we present our ablation results focusing on the feature refinement network, detailed in Sec. 3.1. This network is pivotal for enhancing the consistency of the feature field. The visualization illustrates the energy field utilized for end-effector pose optimization, both with and without the implementation of the refinement network. It is evident that the refined features lead to a concentration of low energy values at positions aligning with the provided hand demonstration. In contrast, the energy field derived from the model without refinement is more dispersed and lacks focused low-energy concentrations.
Ablations on point pruning
Fig. 8 shows our ablation results on the point pruning module (Sec. 3.2). This module not only prunes the outlier points but also increases feature consistency in local neighborhoods, resulting in the improvement of optimization stability and final results of the end-effector poses.
Conclusions
In this study, we address one-shot learning of dexterous manipulations using semantic correspondences from pre-trained vision models. We develop a method to distill features from sparse-view RGBD observations into a consistent 3D field, creating an energy function for optimizing end-effector parameters across varied scenes. Our approach, applied to a dexterous hand, demonstrates robust generalization to new object poses, deformations, geometries, and categories in real-world scenarios.
Given our method’s reliance on a 3D distilled feature field derived from multiview images, it inherits the limitations of rendering-based implicit fields. Specifically, features maintain their meaningfulness only in proximity to object surfaces and begin to lose semantic interpretability as the distance from the surface increases—a phenomenon widely noted in prior works (Kobayashi et al., 2022; Kerr et al., 2023). As a result, our approach excels in learning manipulations involving substantial hand-object contact areas where all hand regions remain close to the object. However, it may struggle in scenarios where certain hand regions are distant from the object. While the method is adept at handling most physical interactions, certain action ”styles” that involve less contact would pose challenges to our approach.
Acknowledgments
The authors would like to thank Prof. Yaodong Yang (PKU) and Prof. Yizhou Wang (PKU) for providing Shadow Hand hardware, and NVIDIA for their generous support of GPUs and hardware. C. Deng, Y. You, and L. Guibas are supported in part by the Toyota Research Institute University 2.0 Program and a Vannevar Bush Faculty Fellowship. Q. Wang, H. Dong, and Y. Zhu are supported in part by the Beijing Municipal Science & Technology Commission (Z221100003422004). Y. Zhu is supported by part by the National Key R&D Program of China (2022ZD0114900), the Beijing Nova Program, and the National Comprehensive Experimental Base for Governance of Intelligent Society, Wuhan East Lake High-Tech Development Zone.