AnyTeleop: A General Vision-Based Dexterous Robot Arm-Hand Teleoperation System

Yuzhe Qin, Wei Yang, Binghao Huang, Karl Van Wyk, Hao Su, Xiaolong Wang, Yu-Wei Chao, Dieter Fox

I Introduction

A grand goal of robotics is to endow robots with human-level intelligence to physically interact with the environment. Teleoperation , as a direct means to acquire human demonstrations for teaching robots, has been a powerful paradigm to approach this goal . Compared to gripper-based manipulators, teleoperating dexterous hand-arm systems poses unprecedented challenges and often requires specialized apparatus that comes with high costs and setup efforts, such as Virtual Reality (VR) devices , wearable gloves , handheld controller , haptic sensors , or motion capture trackers . Fortunately, recent developments in vision-based teleoperation have provided a low-cost and more generalizable alternative for teleoperating dexterous robot systems.

Despite the progress, the current paradigm of vision-based teleoperation systems still falls short when it comes to scaling up data collection for robot teaching. First, prior systems are often designed and engineered towards a particular robot model or deployment environment. For example, some systems rely on vision-based hand tracking models trained on datasets collected in the deployed studio , and some rely on human-robot retargeting models or collision avoidance models trained for the particular robot at use. These systems will scale poorly as the pool of robot models expands and the variety of operating environments increases. Second, each system is created and coupled with one specific “reality”, either only in the real world or with a particular choice of simulators. For example, the HAPTIX motion capture system is only developed for teleoperation in MuJoCo-based environments . To facilitate large-scale data collection with simulation as well as closing sim-to-real gaps, we need teleoperation systems to operate both in virtual (with arbitrary choices of simulators) and in the real world. Finally, existing teleoperation systems are often tailored for single-operator and single-robot settings. To teach robots how to collaborate with other robot agents as well as with human agents, a teleoperation system should be designed to support multiple pilot-robot partners where the robots can physically interact with each other in a shared environment.

In this paper, we aim to set the foundation for scaling up data collection with vision-based dexterous teleoperation, by filling in the aforementioned gaps. To this end, we propose AnyTeleop, a unified and general teleoperation system (Fig. AnyTeleop: A General Vision-Based Dexterous Robot Arm-Hand Teleoperation System), which can be used for:

Diverse robot arm and dexterous hand models;

Diverse realities, i.e. different choices of simulators or the real world;

Teleoperation from diverse geographic locations, via a browser-based web visualizer developed for remote visual feedback;

Diverse camera configurations, e.g. RGB camera with or without depth, single or multiple cameras;

Diverse operator-robot partnerships, e.g. two operators separately piloting two robots to collaboratively solve a manipulation task.

To achieve this goal, we first develop a general and high-performance motion retargeting library to translate human motion to robot motion in real time without learned models. Our collision avoidance module is also learning-free and powered by CUDA-based geometry queries. They can adapt to new robots given only the kinematic model, i.e., URDF files. Second, we develop a web-based viewer compatible with standard browsers, to achieve simulator-agnostic visualization and enable remote teleoperation across the internet. Third, we define a general software interface for visual-based teleoperation, which standardizes and decouples each module inside the teleoperation system. It enables smooth deployment on different simulators or real hardware.

While being very general to support many settings with a single system, our system can still achieve great performance in the experiments. For real-world teleoperation, AnyTeleop can outperform a previous system designed for specific robot hardware with higher success rates on 8 out of 10 tasks proposed in their paper, using the same robot as . For simulated environment teleportation, the smoother and collision-free demonstrations collected by AnyTeleop can bring better imitation learning results with higher success rates on 5 out of 6 tasks proposed in their paper, compared with a previous system specifically designed for that simulator. Finally, we demonstrate that AnyTeleop can be extended to support collaborative manipulation, which to our best knowledge has neither been achieved in the literature of vision-based teleoperation nor on dexterous hands.

Our system is also packaged to be easily-deployable. The containerized design makes installation easy and frees users from handling software dependencies. We are committed to open-sourcing the system and benefiting the community.

II Related Work

Vision-based Robot Teleoperation. Recent years have witnessed an increasing interest in teleoperation of dexterous robot hands by human hands. It relies on accurate tracking of human hand motions and finger articulations. Compared to the costly wearable hand tracking solutions, such as gloves , marker-based motion capture systems , inertia sensors or VR headsets , vision-based hand tracking is particularly favorable due to its low cost and low intrusion to the human operator. Early research in vision-based teleoperation focused on improving the performance of hand tracking and mapping human hand pose to robot hand pose . Recent works have expanded the scope of teleoperating a single robot hand to complete arm-hand systems . However, these systems are designed and engineered towards a particular robot model (e.g., Kuka arm with Allegro hand in , and PR2 arm with Shadow hand in ), and rely on retargeting or collision detection models trained for specific robot hardware (e.g., Allegro hand and XArm6) , making them difficult to transfer to new arm-hand systems and new environments. In contrast, our system is highly modularized with a versatile hand-tracking solution compatible with an arbitrary number of cameras, and configurable robot hand retargeting and motion generation modules for easy adaption to various robot arms and robot hand choices. This allows our system to achieve better performance compared to prior systems on various tasks while generalizing to a set of robot arm-hand systems and multiple environments.

Teleoperation in Different Reality. Manipulation with a dexterous robot hand is challenging due to its high degree of freedom. In recent years, dexterous robot teleoperation has been actively studied and shown promising progress in controlling a multi-fingered hand to perform manipulation tasks in the real world by leveraging the morphological similarity between the dexterous hand and the human hand.

With the advancement of data-driven approaches for robot manipulation , there is a growing need to collect human demonstrations in robotics. To enable easy and scalable data collection, teleoperation has also gained attention in simulated environments . This provides a scalable solution to data collection by eliminating the need for real hardware, while maintaining access to oracle world information. For example, Mandlekar et al. developed a crowd-sourcing platform to teleoperate robots via mobile devices as controllers. Tung et al. further extended this framework to allow multi-arm collaborative teleoperation. The above frameworks rely on inertial sensors for control signals and thus are limited to parallel-gripper and simple tasks such as pick-and-place. Our system offers the ability to perform a wide range of dexterous tasks with robots of different morphologies by utilizing state-of-the-art techniques in perception, optimization, and control. In addition, AnyTeleop is designed to support teleoperation in both virtual and the real world with a unified framework.

III System Overview

Fig. 1 illustrates our proposed paradigms of vision-based teleoperation systems. Below we introduce the features and designs of our system which realize the paradigms.

Any arm-hand. As shown in Fig. AnyTeleop: A General Vision-Based Dexterous Robot Arm-Hand Teleoperation System, AnyTeleop is designed for arbitrary dexterous arm-hand systems that are not limited to any specific robot type.

Any reality. AnyTeleop is decoupled from specific hardware drivers or physics simulators. It can support different realities as visualized in Fig. AnyTeleop: A General Vision-Based Dexterous Robot Arm-Hand Teleoperation System.

Anywhere remote teleoperation. AnyTeleop provides a web-based visualizer to monitor the teleoperation and simulation in standard web browsers, e.g. Chrome.

Any camera configuration. AnyTeleop can consume data from both RGB and RGB-D cameras, and from either single or multiple cameras. Most importantly, it does not require extrinsic calibration as in most previous systems. This allows more flexible camera configurations and lower deployment overhead.

Any number of operator-robot partnerships. AnyTeleop supports collaborative settings where operators separately pilot two robots to collaboratively solve a manipulation task.

Simple deployment. AnyTeleop and all libraries are encapsulated as a docker image that can be downloaded and deployed on any Linux machine, which frees users from handling troublesome dependencies.

Table I compares AnyTeleop with other vision-based dexterous teleoperation systems. We compare the systems in three dimensions: (i) sensor requirements; (ii) robot-related support; (iii) afforded use cases. Among all these teleoperation systems, AnyTeleop is the only one which can support different robot arms and enable collaborative teleoperation. It is also one of the only two systems that can support different dexterous hands.

III-B System Design

The architecture of the teleoperation system is shown in Fig. 2. The teleoperation server (Section IV) receives the camera stream from the driver, detects the hand pose, and then converts it to joint control commands. The client receives these commands via network communication and uses them to control a simulated or real robot. The system is designed with three key principles: modularity, communication-focused, and containerization. Modularity is achieved by implementing well-defined input-output interfaces for each sub-component, allowing for wide applicability to different robot arms, dexterous hands, cameras, and realities. Communication-focused design allows for remote and collaborative teleoperation and reduces computation requirements on the operator’s side by deploying heavy computations on a powerful server. Finally, the containerized design makes installation and deployment easier compared to other robotics systems with heavy software dependencies.

IV Teleoperation Server

The teleoperation server, outlined in Section III, utilizes the RGB or RGB-D data from one or multiple cameras and generates smooth and collision-free control commands for the robot arm and dexterous hand. It consists of four modules: (i) the hand pose detection module, which predicts hand wrist and finger poses from the camera stream, (ii) the detection fusion module, which integrates the results from multiple cameras, (iii) the hand pose retargeting module, which maps human hand poses to the dexterous robot hand, and (iv) the motion generation module, which produces high-frequency control signals for the robot arm. A standardized software interface is defined for all four modules to facilitate flexibility and generalizability in AnyTeleop .

The hand pose detection module offers a unique feature to utilize input from various camera configurations, including RGB or RGB-D cameras, and single or multiple cameras. The design principle is to leverage more information, such as depth, and additional cameras, to improve performance when available. But it can also perform the task with minimal input, i.e. a single RGB camera. The detection module has two outputs: local finger keypoint positions in the wrist frame and global 6D wrist pose in the camera frame. The finger keypoint detection only requires RGB data while the wrist pose detection can optionally use depth information to achieve better results.

Finger Keypoint Detection. Our finger keypoint detection utilizes MediaPipe , a lightweight, RGB-based hand detection tool that can operate in real-time on a CPU. The MediaPipe detector can accurately locate 3D keypoints of 21 hand-knuckle coordinates in the wrist frame and 2D keypoints on the image.

Wrist Pose Detection from RGB-D. We use the pixel positions of the detected keypoints to retrieve the corresponding depth values from the depth image. Then, utilizing known intrinsic camera parameters, we compute the 3D positions of the keypoints in the camera frame. The alignment of the RGB and depth images is handled by the camera driver. With the 3D keypoint positions in both the local wrist frame and global camera frame, we can estimate the wrist pose using the Perspective-n-Point (PnP) algorithm.

Wrist Pose Detection from RGB only. The orientation of the hand can be computed analytically from the local positions of the detected keypoints. However, determining the wrist position in the camera frame can be challenging without explicit 3D information. To enhance MediaPipe for global wrist pose estimation, we adopt the approach used in FrankMocap by incorporating an additional neural network that predicts the weak perspective transformation scale of the hand. The weak perspective transformation approximates the original perspective camera model by assuming that the observed object is farther from the camera than its size. Together with intrinsic parameters, this scale factor can be used to approximate the 3D position of the hand. The wrist position computed this way has a larger error than depth camera, but it is still sufficient for many downstream teleoperation tasks.

IV-B Detection Fusion

The detection fusion module integrates multiple camera detection results. Self-occlusion can be a problem when performing hand pose detection, especially when the hand is perpendicular to the camera plane. Using multiple cameras can alleviate this problem by providing additional views. However, there are two main challenges in fusing multiple detection results: (i) each camera can only estimate the hand pose in its own frame and (ii) there is no straightforward metric to quantify the confidence of each detection result.

To overcome the first challenge, we perform an auto-calibration process using the human hand as a natural marker. We use the first NN frames of hand detection results from multiple cameras to calculate the relative rotation between each camera, expressed in SO(3)SO(3). We find that although the absolute position of detected hand pose is not so accurate in RGB-only setting, the relative motion between consecutive frames is more robust. With orientation between each camera, we can transform the detected relative motion from different cameras into a single frame.

To address the second challenge, we use the SMPL-X hand shape parameters predicted from the detection module, as inspired by Qin et al. . During teleoperation, the true shape parameters should remain constant for a given operator, but the predicted values can contain errors during self-occlusion. We observe that larger shape parameter prediction errors often correspond to larger pose errors. To approximate the confidence score, we take the mean of the estimated shape parameters in the first NN frames as a reference and compute the error between the predicted shape parameters and the reference. Implementation-wise, we require the operator to spread their fingers during the first NN frames to ensure an accurate reference value of shape parameters. The fusion module then selects the relative motion captured by the camera with the highest confidence score and forwards it to the next module. In implementation, we choose N=50N=50.

IV-C Hand Pose Retargeting

The hand pose retargeting module maps the human hand pose data obtained from perception algorithms into joint positions of the teleoperated robot hand. This process is often formulated as an optimization problem , where the difference between the keypoint vectors of the human and robot hand is minimized. The optimization can be defined as follows:

where qtq_{t} represents the joint positions of the robot hand at time step tt, vtiv^{i}_{t} is the ii-th keypoint vector for human hand computed from the detected finger keypoints, fi(qt)f_{i}(q_{t}) is the ii-th forward kinematics function which takes the robot hand joint positions qtq_{t} as input and computes the ii-th keypoint vector for the robot hand, qlq_{l} and quq_{u} are the lower and upper limits of the joint position, α\alpha is a scaling factor to account for hand size difference. An additional penalty term with weight β\beta is included to improve temporal smoothness. When retargeting to a different morphology, such as a Dclaw in Figure AnyTeleop: A General Vision-Based Dexterous Robot Arm-Hand Teleoperation System, we need to specify the keypoint vectors mapping between the robot and human fingers manually. It is worth noting that this module only considers the robot hand.

IV-D Motion Generation

Given the detected wrist and hand pose, our goal is to generate smooth and collision-free motion of robot arm to reach the target Cartesian end-effector pose. Real-time motion generation methods are required to have a smooth teleoperation experience. In the prior work of , the robot motion is driven by Riemannian Motion Policies (RMPs) that can calculate acceleration fields in real-time. However, accelerations towards a particular end-effector pose do not guarantee natural trajectories. In this work, we adopt CuRobo , a highly parallelized collision-free robot motion generation library accelerated by GPUs, to generate natural and reactive robot motion in real-time. In AnyTeleop, the motion generation module receives the Cartesian pose of the end-effector at a low frequency (2525 Hz) from the hand detection and retargeting modules, and generates collision-free joint-space trajectories within joint limits at a higher frequency (120120 Hz). The generated trajectories are ready for safe execution by impedance controllers on either a simulated or real robot.

V Web-based Teleoperation Viewer

To better support the teleoperation tasks, we implement a web-based visualization module to facilitate remote and collaborative teleoperation, especially for teleoperation in simulated environments. It has the following features: (i) browser-based viewer, which makes it easily accessible remotely; (ii) synchronized visualization, i.e. two operators working on the same collaborative task should see the same scene synchronously from their own local view ports. The viewer is developed based upon the meshcat library and utilize Three.js for rendering. The visualization server ports the simulation results onto the browser after each simulation iteration. Operators can get visual feedback from the browser window and move their hands to control the corresponding robot. More details about the implementation of our viewer can be found in the supplementary materials.

VI System Evaluation

We perform profiling on modules mentioned in Section IV on a desktop and a laptop. As shown in Table II, the most time-consuming module is hand pose detection, which runs on a GPU for real-time inference. The designed maximum frequency for hand pose detection is 25Hz, so both the desktop and laptop can meet the requirement. Both the retargeting module and the fusion module run at the same frequency as the hand detection module due to the publisher and subscriber logic. For best performance, the motion generation module should run at 120Hz but can still work with a lower frequency. Notably, we found it difficult to achieve this throughput when running all these modules on the same computer. Luckily, with our communication-oriented design, we can run the control modules on a separate machine to achieve the best performance.

VI-B Real Robot Teleoperation

In this section, we will test our AnyTeleop system across a wide range of real-world tasks that covers diverse objects and manipulation skills. Besides, we will compare our teleoperation performance of AnyTeleop with a similar teleoperation system. A fair comparison of real-robot tasks is often very challenging due to the difficulty in replicating the baseline methods delicately. To ensure a more fair comparison, we replicate the ten manipulation tasks proposed in Robotic Telekinesis with the same XArm6 robot, Allegro hand, and similar objects. A trained operator attempt to solve this tasks using AnyTeleop system. The ten tasks are visualized in Fig. 3. Same as , we run each task ten times for AnyTeleop and use a single Intel RealSense camera. For the baseline method, we directly use the results reported in their paper.

As shown in Table III, AnyTeleop can get a higher success rate of 8/10 tasks and the same success rate on 2/10 compared with the baseline. Although AnyTeleop is designed to be more general, it can still outperform the baseline system that was specifically designed for the XArm6-Allegro hardware. We find that the major advantage of our system is the capability to handle objects with thin-walled structures, such as the cup-stack, two-cup-stacking, and cup-into-plate tasks. Our optimization-based retargeting module can close the distance between finger tips, which makes grasping the cup more stable. However, the network-based retargeting can hardly translate the fine-grained precision grasp from human to robot, which leads to a lower success rate.

VII Applications

The most important application of the proposed system is imitation learning from demonstration. We can first collect demonstrations on several dexterous manipulation tasks and then use the data to train imitation learning algorithms. In this experiment, we will show that the teleoperation data collected using our AnyTeleop can better support downstream imitation learning tasks. In the following subsection, we will first introduce the experiment setting and baseline and then discuss the experimental results.

Baseline and Comparison. To fairly compare with previous teleoperation systems, we need to align both the task setting and robot configuration precisely. It is often challenging for real-robot hardware but much easier for a simulated environment. Thus, we choose a recent vision-based teleoperation work that can be used for simulated robots as our baseline. It is worth noting that we are comparing two teleoperation systems via the demonstration data collected by each system. Thus, we compared with the baseline by training the same learning algorithm on different demonstration data collected via the baseline system and our teleoperation system. We follow to choose Demo Augmented Policy Gradient (DAPG) as the imitation algorithm. We also compare it with a pure reinforcement learning (RL) based algorithm from which does not utilize demonstrations. We provide the same dense reward for RL training as previous work .

Manipulation Tasks. We directly use the manipulation tasks proposed by the baseline work for comparison, which include three tasks: (i) Relocate, where the robot picks an object on the table and moves it to the target position; (ii) Flip Mug, where the robot needs to rotate the mug for 90 degrees to flip it back; (iii) Open Door, where the robot needs to first rotate the lever to unlock the door, and then pull it to open the door. These tasks are visualized in the left figure of Table VI. The manipulated objects in all three tasks are randomly initialized and the target position is also randomized in Relocate. Each manipulation task has two variants: the floating-hand variant and the arm-hand variant. The floating-hand is a dexterous hand without a robot arm that can move freely in space. The arm-hand means the hand is mounted on a robot arm with a fixed base, which is a more realistic setting.

Demonstration Details. For the baseline teleoperation system , we directly use the demonstration collected by the original authors with 50 demonstration trajectories for each task. The baseline system only utilizes a single RGB-D camera. For fairness, we also collect 50 trajectories for each task using the single camera setup. The baseline system can only handle floating hands and they propose a demonstration translation pipeline to convert the demonstration with floating hands to demonstrations with arm-hand. For our AnyTeleop , we collect demonstrations using the arm-hand setting and convert the demonstration to floating hand so that the demonstration can be used by both the floating-hand variant and arm-hand variant.

Results and Discussion. For each method on each task, we train policies with three different random seeds. For each policy, we evaluate it on 100 trials. More details about the success metrics can be found in . As shown in Table VI, the imitation learning algorithm trained on demonstration collected by AnyTeleop can outperform baseline and RL on most tasks with one exception. Compared with the demonstration collected via the baseline system, our system has two benefits that contribute to better performance in imitation learning: (i) The collected trajectory is more smooth, which means that the state-action pairs are more consistent and easier to be consumed by the network. (ii) Different from the baseline, our system explicitly supports teleoperation with arm-hand system and guarantees no self-collision. On the contrary, the baseline system utilizes retargeting to generate joint trajectory for robot arm, which may lead to several self-collision for robot arm. Thus we can observe significant performance gain of our system for manipulation tasks with arm-hand. For the flip mug task, the difficulty of collecting demonstration with arm is much larger than with a floating hand, which influences the demonstration quality.

VII-B Collaborative Manipulation

Collaborative manipulation is a key technology for the development of human-robot systems . Collecting demonstration data for collaborative manipulation tasks has been a challenging task since it requires multiple operators to work together seamlessly. With our modularized and extensible system design and web-based visualization, our system enables convenient data collection on collaborative tasks, even if operators are not in the same physical location. In this section, we show that our teleoperation system can be extended to a collaborative setting where multiple operators coordinate together to perform manipulation tasks. We choose human-to-robot handover as an example as shown in Fig. 4. In this setting, operator #1\#1 control a robot hand, and operator #2\#2 control a human hand.

Collaborative Teleoperation System Design. Fig. 5 illustrates the system architecture for multi-operator collaboration, which includes two components. (i) Teleoperation Units: It is composed of a computer that is connected to at least one camera and a human operator. In each teleoperation unit, the human operator will watch the real-time visualization on a web browser and move the hand accordingly to perform manipulation tasks. (ii) Central Server: it runs the physical simulation and the web visualization server. The detection results from multiple teleoperation units are sent to the server and converted into robot control commands based on the pipeline in Section IV. Meanwhile, the web visualizer server will keep synchronized with the simulated environment and maintain the visualization resources as described in Section V.

VIII Failure Modes

As illustrated on the project page, we have identified two failure modes: (i) loss of tracking during fast human hand motion, which triggers a pause and re-detection process; (ii) unreliable hand pose when the hand is in self-occlusion. For mode (i), the workaround is to instruct the operator to slow down their hand motion. The issue (ii) can be solved by incorporating multiple cameras for tasks that require significant hand rotation.

IX Conclusion

In this paper, we introduced AnyTeleop, a versatile teleoperation system that can be applied to diverse robots, assorted reality, and varied camera setup, and can be operated by a flexible number of users from any geographic locations. The experiments show that AnyTeleop outperforms previous systems in both simulation and real-world scenarios while offering superior generalizability and flexibility. Our commitment to an open-source approach will facilitate further research in the field of teleoperation.

X Acknowledgement

We express our gratitude to Ankur Handa, Balakumar Sundaralingam, and Nick Walker for their insightful discussions throughout the development process of the motion control module in AnyTeleop. We would also like to extend our thanks to Isabella Liu, An-Chieh Cheng, Ruihan Yang, Yang Fu, Linghao Chen, Jiarui Xu, Xinyu Zhang, Xinyue Wei, Jiteng Mu, and Jianglong Ye for their efforts in testing and evaluating the teleoperation system.

References

-A Supplementary Overview

This supplementary material provides more details, results and visualizations accompanying the main paper, including

More details and visualization about teleoperation server, including detection and retargeting modules;

More details about web-based teleoperation viewer;

Additional experimental results on system evaluation.

More visualization can be found at our project page: http://anyteleop.com.

-B Teleoperation Server

In this section, we will show more intermediate results from our system, including visualization of hand pose detection results and the retargeting results of various robot hands.

Visualization of Hand Pose Detection We visualize the hand pose detection results in Figure 6. We showcase five typical cases, which include: (i) a hand spreading out the fingers for teleoperation initialization, (ii) fingers facing downwards in preparation for a top-down grasp, (iii) a precision grasp using the thumb and index finger, (iv) a power grasp using all five fingers, and (v) a failure case where the hand is positioned vertically relative to the camera plane.

Visualization of Hand Pose Retargeting We demonstrate the results of hand pose retargeting in Figure 9. The figure displays seven gestures being performed using four different dexterous hands.

-C Web-based Teleoperation Viewer

In this section, we demonstrate how the web-based visualizer provides accessibility and multi-view support for teleoperation through its lightweight rendering and capability to run in multiple browser windows. Figure 7 shows screenshots of the web-based visualizer when it is used to visualize the five IsaacGym tasks depicted in the Figure 1 in the main paper.

Lightweight Rendering vs High Visual Quality. The design of our web-based viewer prioritizes accessibility and convenience, as it can be used on any device with a browser and provides minimal but sufficient rendering capabilities for teleoperation. Although the rendering quality may not be as advanced as the original simulator viewer, simulation states can be saved for offline rendering to produce high-quality visual data. For example, in visual reinforcement learning tasks using RGB images as inputs, the rendered data can be generated using a more powerful engine such as a ray tracer after teleoperation is completed.

Multi-View Support for Teleoperation. In teleoperation, human operators often require a clear understanding of the spatial relationships between objects and robots to make informed decisions. This information can be provided through multi-view rendering, which is a widely used technique in previous teleoperation works . Our web-based viewer offers multi-view support to the operator by simply opening multiple browser windows. As shown in Figure 8, an example of the operator using two views to perform a manipulation task is displayed. The operator is able to open as many windows as needed to enhance their teleoperation experience.

Appendix A System Evaluation on Camera Configurations

In this section, we examine the impact of different camera configurations on the teleoperation performance of our system, AnyTeleop , which is capable of supporting diverse configurations including RGB, RGB-D, and single or multiple cameras. Even with a minimal configuration, i.e. a single RGB camera, the system can still perform effectively. Additionally, by adding more resources, such as multiple cameras, our system can achieve better performance.

We use the Play Piano task implemented in IsaacGym as the evaluation scenario, which requires the robot hand to press piano keys in a specific order. The task is shown in the Figure 1 in the main paper and the video in the supplementary material. To quantify performance, we introduce two task metrics: (i) completion time, i.e. the elapsed time from start to finish, and (ii) the percentage of incorrect key presses, which measures the number of incorrectly pressed keys relative to the total number of keys.

A trained operator performs the task ten times for each camera configuration. As reported in Table IV, with additional information, such as depth, and increasing number of cameras, the task can be completed faster and with fewer errors, which demonstrates that our system allows users to easily trade-off between efficiency and system cost based on their use case.