RoboTAP: Tracking Arbitrary Points for Few-Shot Visual Imitation

Mel Vecerik, Carl Doersch, Yi Yang, Todor Davchev, Yusuf Aytar, Guangyao Zhou, Raia Hadsell, Lourdes Agapito, Jon Scholz

I INTRODUCTION

Imagine if you could track an arbitrary number of points in space, in any scene, through occlusions, motion, and deformation – how might it simplify the manipulation problem? Recent advancements in dense-tracking have provided exactly this capability, but it has yet to be explored in a manipulation setting. In this paper we investigate dense-tracking as a perceptual primitive for manipulation, with the goal of producing a single easy-to-use system that can solve a wide range of problems without task-specific engineering.

This objective is closely related to recent work on foundation-models for robotics , which treats manipulation across embodiments and tasks as a large-scale sequence-modeling problem. These approaches are incredibly general and powerful, especially when utilizing models pre-trained on large multi-modal datasets , but they tend to be data-hungry and expensive to train. From this perspective, our primary interest is in understanding the representational and computational bottlenecks in the manipulation problem, in order to improve the scalability and performance of generalist robotic systems.

Our hypothesis is that much of the complexity in low-level manipulation can be reduced to three fundamental operations: (1) identifying what is relevant in the current frame, (2) identifying where it is, and (3) identifying how to move it in desired directions. We show that all three operations can be parameterized via dense-tracking, yielding a general formulation for manipulation that does not require task-specific engineering. Furthermore, we find that point tracks can provide an interface between these different operations, allowing us to factorize the problem into simple low-dimensional functions which, together, can express a large space of possible behaviors.

The behaviors of interest span a range of real-world problems involving precise multi-object rearrangement, e.g. pick-and-place, (high-clearance) insertion, and stacking. Our approach, RoboTAP, allows few-shot imitation learning of these behaviors in minutes, which we highlight via a task that involves grasping a glue-stick, applying it to a surface, mating the two parts, and placing the result at a desired location. This task requires over 1000 precise real-valued actions to be successful, and is messy and irreversible, which renders it infeasible for approaches requiring large-scale data-gathering.

For all of these tasks, RoboTAP automatically extracts the individual motions, the relevant points for each motion, goal locations for those points, and a generates a plan that can be tracked by a low-level visual servoing primitive. While currently less general than fully end-to-end approaches, our approach can be trained on as few as 4-6 demonstrations per task, does not require action-supervision, and effortlessly generalizes across clutter and object pose randomization. Our main contributions are as follows: (1) RoboTAP, a formulation of the multi-task manipulation problem in terms of dense-tracking, (2) Concrete implementations of RoboTAP’s what, where, and how problems in the form of visual-saliency, temporal-alignment, and visual-servoing, (3) A new dense-tracking dataset with ground-truth human annotations tailored for our tasks and evaluated on the TAP-Vid Benchmark focusing on Real-World Robotics Manipulation, and (4) Empirical results that characterize the success and failure modes of RoboTAP on a range of manipulation tasks involving precise multi-body rearrangement, deformable objects, and irreversible-actions.

II Related work

Regressing actions directly from images in an end-to-end manner is one established way for utilising visual inputs . While in theory this allows for creation of arbitrarily complex policies in practice this often requires a large amount of data including the target objects or classes.

II-A2 Pose-based Approaches

Another approach is to define a policy on top of object poses regressed from images . Pose is a strong signal for object state, but is not well-defined for deformable or symmetric objects, and is difficult and time-consuming to supervise accurately and generally.

II-A3 Keypoint-based Approaches

Within RoboTAP, point locations are detected from an RGB camera and converted to arm motions in a visual feedback-loop, which is commonly referred to as image-based visual servoing . Existing autonomous visual servoing techniques typically rely on hand-designed target detectors or confidence maps , which lack the generality required for few-shot imitation in arbitrary scenes.

A closely-related approach is to define a small sparse set of keypoints for a given object class Choosing a different point set however requires a retraining of the whole model. Methods such as TACK or DON generalize this approach by learning an embedding for any point on an object in the image. This still falls short of representing arbitrary points in scenes, and must be retrained in order to adapt to new class of objects.

As a result, keypoint-based approaches can perform very well in specific settings, but act as an extra barrier to deploy robotic systems on novel tasks. In this work, we aim to demonstrate that recent advances on point tracking models such as are good enough to act as a sole perception model for both motion tracking, cross scene correspondence and segmentation. This system doesn’t use any depth information during training or evaluation and uses only a single non-calibrated gripper mounted camera.

Lastly, the “what” / “where” factorization we explore in this paper has several precedents in the robotics literature, e.g. , and has been more broadly explored in neuroscience .

II-B Point Tracking in Computer Vision

Within computer vision there have been approaches which extract points from any scene without any pretraining. For example non-learned descriptors such as SIFT or ORB require no class specific training, but they produce many spurious matches which cannot be easily filtered on non-rigid scenes, making them difficult to use as a target for visual servoing. Recently there have been several advances in high-performance long-horizon visual tracking such as OmniMotion , PIPs , TAP-Net or TAPIR , which provide high-quality tracks with very few outliers. This means that Euclidean distance in point space, with no outlier removal, can serve as a reliable metric of whether two spatial configurations are similar, or whether two motions are similar. While all of these approaches are made to run on whole videos, we note that the setup (especially TAPIR) can be modified to run online on a frame-by-frame basis, making it suitable for robotics applications.

III Approach

The quantities qtq_{t} and gtg_{t} represent the “what” and “where”, respectively, of the current action, and are extracted from DD via a procedure gt,qt=f(st,D)g_{t},q_{t}=f(s_{t},D), described in Section III-A. For simplicity, in this paper we don’t implement ff as a per-step operation, but rather as a one-time plan that summarizes the objectives common across DD. An overview of this procedure and it’s connection to the control is illustrated in Fig. 2. The high-level outline for our overall solution is as follows:

Sample large number of query points from DD and track those points across all trajectories using TAPIR.

Segment the trajectories into phases, and discover a descriptor-set qtq_{t} for each phase which characterize the motion.

Pack the sequence of qtq_{t} and corresponding trajectory slices (representing the goals) into a “motion-plan”.

Execute this motion plan using the low-level controller in a sequential fashion, advancing stages based on a simple final-error criterion.

The first core step of the RoboTAP approach, is to extract important motion from demonstrations in order to construct a motion-plan. This seen in Fig. 2. In order to accomplish this we assume that our demonstrations DD are comprised of an unknown but equal number of phases, each involving the motion of a particular set of points. Although the set of points could be learned end-to-end (e.g. by behavioral cloning), this would require more data than we have available. Therefore, we instead explicitly identify motion segments that are shared across demos, and then discover a set of points for each that are moving consistently, which we term active points.

Temporal alignment has been well-studied , but for simplicity, we rely on the assumption that our tasks consist of grasping and releasing objects or making contacts, which means that temporal segments are trivially obtained by thresholding gripper actions and forces.

Active Point Selection. Given a temporal segment, identifying active-points involves asking what points qt∈Qq_{t}\in Q the low-level controller πl\pi_{l} should move in order to generate motions that accomplish a similar result as observed in the segment. Our criteria for point selection are illustrated in Fig. 3. The key insight is that the active points may begin at diverse locations, but for goal-directed behavior they tend to end up the same place across all demos at the end of the relevant segment. Therefore, we select points which end at the same place at the end of the demo segment (relative to the camera), and remove points which don’t move at all across the segment (i.e., the gripper and any already grasped object).

While this alone can identify many of the desired active points, not all of the useful points are guaranteed to be visible at the end of each motion segment, either due to occlusion or simply failures by the detector; furthermore, there may be inconsistencies due to imprecise demos that make it difficult to set thresholds on whether a point is moving. Therefore, we recast the active point discovery as votes on which object (or object-part) is being manipulated, which we substantially improves robustness. However, this means we must perform unsupervised object discovery from the demos, a classically difficult problem which we find is rendered surprisingly reliable given robust dense point tracks. We then combine these heuristics with a voting-based scheme, whereby we select the clusters which contains the most points which are “active” according to the first two criteria.

Clustering. There are numerous approaches to object-based clustering, ranging from semantic segmentation to generative modeling , but for simplicity, we use motion estimates extracted from TAPIR, since this means we do not require semantic labels, and we find it can work reliably from a remarkably small amount of data.

The simplicity of this equation is somewhat remarkable, as it is essentially standard bundle adjustment , but we do not model outliers. Outlier rejection is a critical component of prior structure-from-motion methods as they are typically based on descriptors like SIFT , which have a high proportion of extreme errors. This severely limits their ability to do multi-object reconstruction, as it is difficult to distinguish between small objects and outliers. TAPIR, however, has powerful occlusion estimation, meaning that all predicted points should be modeled. We informally tried approaches which used robust losses such as L1L1, Huber, and even truncated L2L2 losses with appropriate RANSAC-style proposals, but did not find formulations that performed on par with simple L2L2.

We parameterize both Pi,kP_{i,k} and At,kA_{t,k} using neural networks, which aim to capture the inductive biases that points nearby in 2D space, and also frames nearby in time, should have similar 3D configurations. Specifically, Pi,k=P(pi,oi∣θ1)kP_{i,k}=\boldsymbol{P}(p_{i},o_{i}|\theta_{1})_{k}, where θ1\theta_{1} parameterizes the neural network P\boldsymbol{P} which outputs a k×3k\times 3 matrix, and At,k=A(ϕt∣θ2)kA_{t,k}=\boldsymbol{A}(\phi_{t}|\theta_{2})_{k}, where ϕt\phi_{t} is a temporally-smooth learned descriptor for frame tt, θ2\theta_{2} is a neural network parameter for neural network A\boldsymbol{A} which outputs a k×3×4k\times 3\times 4 tensor representing rigid transforms.

The above optimization is, unsurprisingly, somewhat difficult due to local minima; we find that these local minima can be avoided by splitting clusters during optimization. For details on this and the neural networks, see Appendix Section B. Given a clustering, we select clusters by voting. Every previously selected active point casts a vote for a cluster and we merge clusters with largest number of votes. To further avoid selecting points from the gripper we compute average point movement for each cluster and remove clusters where this movement is below a threshold. Finally, we discard any points are not visible in the current phase, or which are not nearby other points on most frames, as these are likely to be tracking failures. For details, see Appendix Section E.

III-B Robot controller

In theory, to obtain the correct image Jacobian we need to consider the camera intrinsics and the extrinsics. However in our case we only require that it points in the correct direction, and we then tune the rotation and translation gains of the controller for stability. To demonstrate this in all of our experiments we assume vertical field of view of 90∘90^{\circ} and unit depth which leads the following image Jacobian for a single point:

Where the columns correspond to end-effector-frame linear and angular velocity of the gripper, and u, v are normalized image coordinates such that vertical dimension is within . Using this Jacobian, we then compute the gripper motion that would by minimize the L2 error between the current detections pp and goal locations gg under the linear approximation, following standard visual servoing. Because the error is typically dominated by translation, applying the Jacobian naively would result in the controller explaining translation via rotation and scaling; therefore, we use Gram-Schmidt orthogonalization to eliminate the average translation before computing the Jacobian with respect to rotation and scaling.

If TAPIR provided perfect point locations, the above algorithm would work effectively without modification. However, for precise tasks, TAPIR errors can still lead to two specific failure modes. First, outliers from TAPIR can overwhelm the controller when true errors are low. To deal with this, we leverage TAPIR’s uncertainty outputs, and only use points which are confidently predicted for both the current frame and in the demo. Second, noisy detections introduce a statistical bias towards minimizing the spread of the point cloud in order to reduce unexplainable errors. This presents itself as a bias towards moving the camera further back. Therefore, we modify the controller by leveraging the relation between the problems of aligning pp to gg and the inverse alignment of gg to pp. This changes the zz-axis update to perform variance matching rather than directly optimizing the visual servoing objective. To further reduce variance near the end of each servoing stage (when the demo locations are assumed to agree), we use an average of point locations across all demos as a target to further reduce variance, while we use a single demo as a target earlier in each stage. Further details can be found in Appendix Section F.

III-C Online TAPIR

In the previous section we describe how to use TAPIR point tracks inside a robot controller. However one of the hurdles we must overcome to do this, is to find a way to run it online, inside a control-loop. The original TAP-Vid was proposed as an offline benchmark, in that methods may process the entire video before computing a trajectory, and top-perfoming methods like TAPIR , PIPs , OmniMotion , and TAP-Net all process a full video at once. Therefore, a key contribution is to reformulate the top-performing models (i.e. TAPIR) to work online without harming its accuracy.

We first note that TAPIR can be divided into three stages: 1) extracting query features for query points, 2) initializing the track via global search on every frame, and 3) refinement with a temporal ConvNet. To address the first stage, we note that the query features depend only on the query frame, so it is straightforward to extract query feature computation into a separate function. For the second stage, the initialization for a given frame depends only on the query features and the features for that frame, meaning the initialization can be computed one frame at a time.

The third stage, however, is a refinement, which is not straightforward to run online: at its core is a convolutional neural network which runs across time. The solution is to observe that this can be converted into an online model by replacing the temporal convolutions with similarly-shaped causal convolutions , such that the activation of each unit at time tt depends only on the activations from times ≤t\leq t. At training time, this is a trivial architectural change. When running on the robot, however, it means we must keep a history of input activations for each layer with a temporal receptive field. Any activations that a unit at time tt depends on are read from the history. Once the forward pass is complete, a new history is constructed and returned which can be used at the next timestep. In practice, we find that the performance impact of moving to a causal model is minimal (see tables III and IV). See Appendix C for details.

IV EVALUATION

In this section, we aim to demonstrate the strengths and limitations of our system and its components. The capabilities of RoboTAP overall are directly linked to the precision of TAPIR itself on the domain we are interested in. Therefore to evaluate and advance its abilities we introduce a new point tracking dataset focused on robotics settings. Further, since the TAPIR model we are using has been modified to be run in an online way, we evaluate the effect of these changes on it’s performance on the TAP-Vid benchmarks.

Following this, we show evaluations of the full RoboTAP system. In particular, we show that RoboTAP can tackle complex tasks with many stages from just a handful of demos, can deal with start and end poses significantly different from those in the demos, can both place objects precisely and follow trajectories precisely, has strong invariance to distractors in the workspace, and can deal with non-rigid objects.

We choose a variety of tasks involving placement, insertion, and gluing, that are representative of real-world assembly problems. While all of the demonstrations were gathered without occlusions and distractor objects we explore how the ability of RoboTAP is affected when these are present. Finally we also study the servoing precision using a calibrated real-world setup. Further experimental ablations of the controller in simulation can be found in Appendix Section A.

The ability to track and relate points in the scene is what enables RoboTAP to generalize to novel scenes and poses. These capabilities are directly powered by the performance of these models and therefore a core part of our contribution is enabling advances in these areas. To enable better research we introduce an addition to the previous TAP-Vid benchmark in a form of a new dataset which focuses specifically on robotic manipulation.

Specifically we collect 265 real world robotics manipulation videos. The data source mainly comes from teleoperated episdoes in the public DeepMind robotics videos . We annotate each video with human groundtruth point trajectories, following the same instruction in the previous TAP-Vid dataset , with 5 points per object and 5 points on the background. To make the evaluation closer to the realistic manipulation setup, we sampled the videos from different camera viewpoints (i.e. basket view, wrist view) where basket view is static and wrist view moves along with the gripper.

Table I shows the overall statistics of the newly collected RoboTAP dataset. Comparing to the existing simulated TAP-Vid-RGB-Stacking dataset, it contains more videos, more points, and more frames on average, but more promisingly real world. Table II shows more details of the annotated points in different camera setups. Besides the general evaluation, we further provide labels of whether camera is static or moving and a point is static or moving.

IV-B Online TAPIR

We compare our online TAPIR model with existing baselines on both the TAP-Vid benchmark and the new RoboTAP dataset. Results in Table III show that our online TAPIR achieves accuracy similar to the state-of-the-art offline TAPIR model, significantly outperforming TAP-Net and PIPs. On RoboTAP, it achieves an average jaccard (AJ) of 59.1, nearly matching offline TAPIR’s 59.6. Table IV shows more detailed evaluation on different dataset splits.

IV-C Robot Setup

To run our system in the real world we use the Franka Panda Emika robot in the impedance control mode, 2f-85 Robotiq gripper and FT 300 Force Torque Sensor. The control signal was sent at 10Hz and interpreted as velocity of the impedance controller setpoint. To gather images we used a Basler Dart daA1280-54ucm at a resolution of 640x480 pixels mounted to the wrist using a custom 3D printed adapter. Our camera images were undistorted, but we did not use intrinsic or extrinsic calibration of our camera.

To collect the demonstrations we move the robot via the Franka Panda cuff. RoboTAP is compatible with kinesthetic teaching because, unlike much other imitation learning work, it does not require access to actions beyond gripper opening/closing state. During teaching we record robot positions, forces from a wrist force-torque sensor, and wrist camera videos at 10Hz. We used a keyboard to open and close the gripper. It takes about 30 seconds to record a demonstration for a pick and place task, 1 minute for the gluing task and up to 2 minutes for the 4-object insertion task.

We used a total of 9 tasks described in Table V, some of which are illustrated in Figure 4 - see our website for full list of task videos. When running this controller we have noticed a very repeatable pattern of cases where the controller succeeds and where it fails. For a majority of the tasks we have attempted we observe robust performance provided the algorithm’s basic assumptions are met, yet we have seen it fail in a few specific cases. Therefore we believe that a simple success metric would not appropriately capture its behaviour and we aim to provide other avenues to express its performance.

In all cases we aimed to demonstrate the performance of the controller in a clean environment, and then show how performance degrades with increased clutter and occlusions. The only tasks where we have not observed reliable performance were the LEGO stack, in which the controller lacked the precision to stack the bricks, and the 4 object stack, where the compounding of the errors often lead the last object to be dropped in the wrong location. Overall we observed 4 factors which caused failures: (1) Occlusions cause failures when the active points cannot be seen. This most notably happens due to the gripper occlusions. (2) Scale changes cause failures when the gripper moves too far away and none of the currently-tracked features can be matched because the objects are too small. (3) Distractors can cause failures when the scene contains a similar object to the essential object such as a red apple vs a tomato. (4) Collisions can happen with the gripper or currently-held object as the controller does not have a way to perceive clutter. However in absence of these cases, we observed robustness to novel arrangements and even dynamic changes of the scene while the controller is being executed.

IV-D Real world precision

Previous experiments demonstrated the qualitative behavior of several long-horizon manipulation tasks. We conclude by evaluating the precision of our visual-servoing controller quantitatively. For this we take inspiration from position-repeatability analyses performed on commercial industrial robot arms. We gathered 6 demonstrations of putting a textured white plastic gear on a square of graph paper, which allowed us to precisely compute of the placement error using OpenCV . For the demonstrations we used 3 different configurations of the white gear and the target pattern. The target had a cross in the middle which allowed us to take pictures of the centre of the gear and automatically extract its location.

Then we ran the controller 30 times for each of 3 novel goal locations. First goal location, Near, was positioned in the middle between the locations which were used for demonstrations. The second one, Far, was on the opposite side of our workspace and the last one, Rotated, was on the side but rotated by 90 degrees. Notably the second 2 goal positions are outside of the distribution provided during demonstrations.

We can see the results of this evaluation in Fig. 6. Based on the figure we can see that the error of the controller is comparable to the error seen across the demonstrations. The error along the plate, orthogonal to the grasp direction is on the order of several milimeters. Based on these results we can see that the controller in its current form would not be suitable for sub-mm insertion tasks, however we have not attempted to extensively tune it for this purpose. Further information about these experiments is provided in Appendix Section D.

V CONCLUSION

We presented RoboTAP, a manipulation system that can solve novel visuomotor tasks from just a few minutes of robot interaction. RoboTAP does not require any task-specific training or neural-network fine-tuning. Thanks largely to the generality of TAP, we found that adding new tasks (including tuning hyper-parameters) took minutes, which is orders of magnitude faster than any manipulation system we are familiar with. We believe that this capability could be useful for large-scale autonomous data-gathering, and perhaps as a solution real-world tasks in its own right. RoboTAP is most useful in scenarios where quick teaching of visuomotor skills is required, and where it is easy to demonstrate the desired behavior a few times.

There are several important limitations of RoboTAP. First, the low-level controller is purely visual, which precludes complex motion planning or force-control behavior. Second, we currently compute the motion plan once and execute it without re-planning, which could fail if individual behaviors fail or if the environment changes unexpectedly.

Several of the ideas in RoboTAP, e.g. explicit spatial representation and short-horizon visuo-motor control, could also be applicable in more general settings. In the future we would like to explore whether RoboTAP models and insights can be combined with larger-scale end-to-end models to increase their efficiency and interpretability.

ACKNOWLEDGMENT

We would like to thank Tom Rothörl, Akhil Raju, and Marlon Gwira for their help with our robot setup and Dilara Gokay for dataset infrastructure support. We would also like to thank Nicolas Heess, Oleg Sushkov, Andrew Zisserman and João Carreira for useful discussions and feedback.

References

Appendix A Simulated ablations

To explore the choices behind our visual servoing controller and quantitatively evaluate them we define a MuJoCo simulated environment . Within this environment we place a camera and an object from the Google Scanned Objects dataset . We create a dataset of 480 trajectories with 20 different objects and 24 motions per object. There trajectories contain changes up to 5x in scale and movements across the whole image. The task is for the controller to follow a demonstrated trajectory using the point tracks provided by TAPIR. To make the environment more visually complex we add a spatially fixed background to each trajectory which is randomly sampled from the real demonstration data collected for our tasks. This acts as a distraction for the TAPIR model and increases the observed noise to better match the real trajectories.

Our visual servoing controller needs a set of points qq to be tracked. In simulation we construct this by generating an extra video with a different background to the demonstration and randomly sample points across different timesteps of the demonstration. To evaluate our controller we sample a yet different background and match the initial relative camera and object pose to the one seen in the demonstration. Then at each timestep the controller outputs an action which is a desired pose difference to the next state. We run this for each of the 480 trajectories and evaluate in how many of the videos the controller converges to the final state of the trajectory. To evaluate visual servoing with 6 degrees of freedom (DOF) we used the jacobian described in (3) instead of (1). We note that these visual servoing tasks are very challenging as they contain large changes of scale and the tracked object presents only a very small part of the scene. The results of these experiments are presented in Table VI. First we can see that this approach can work when all 6 degrees of freedom are being used for servoing, however it is significantly less stable. If we do not evaluate the jacobian at both demonstration and current location, we also observe a significant drop in performance due to the statistical bias. We see similar but smaller performance drop when we do not orthogonalise the jacobian, i.e. if we replace J⊥J^{\perp} with JJ. Both of these issues are especially notable in cases which contain large changes in scale.

Appendix B Clustering Implementation Details

Where ii indexes points, tt indexes video frames, and kk indexes the clusters. We parameterize PP and AA with simple neural networks. This problem is somewhat underconstrained and difficult to optimize, so we use neural networks to inject some priors.

Here, θ\theta parameterizes the neural networks that output AA and PP, and includes ww and both of the ‘fork’ variables w′w^{\prime} and w′′w^{\prime\prime}. After a few hundred optimization steps, we replace ww with wKw^{K}, and create new ‘forks‘ of this matrix (initializing the forks with small perturbations of wKw^{K}). We begin with k=1k=1 and repeating recursive forking process until the desired number of objects is reached. We find that this recursive splitting can result in over-segmentation, but this can be somewhat mitigated by deleting clusters after optimization is finished. This is an analogous process: we optimize for the minimum loss after clusters are deleted. Our full implementation of this algorithm is available in our public project repository.

Appendix C Implementation details of online tapir

In this section, we provide implementation details for our causal version of the TAPIR model. Our full implementation can be found in our project Github.

Recall that the temporal point refinement of the TAPIR model uses a depthwise convolutional module, where the query point features, the xx and yy positions, occlusion and uncertainty estimates, and score maps for each frame are all concatenated into a single sequence, and the convolutional model outputs an update for the position and occlusion.

The depthwise convolutional model replaces every depthwise layer in the original model with a causal depthwise convolution; therefore, the resulting model has the same number of parameters as the original TAPIR model, with all hidden layers having the same shape. There is no change to the training procedure, as we can concatenate all frames in the training sequences and run the convolutions across the full sequence.

At test time, however, we must preserve a “causal context” across frames, which contains the activations computed for the last frame that are required as input for the depthwise convolution layers. Because the temporal receptive field of the depthwise is 3, we must keep 2 frames of context for every depthwise convolution. The refinement steps use 4 PIPs iterations, each of which has 12 blocks, and each block has 2 temporal depthwise conv layers with 512 and 2048 units respectively. Therefore, the temporal context is 2×(512+2048)×12×4=245K2\times(512+2048)\times 12\times 4=245K floating-point values for every point.

Appendix D Precision experiments - further information

In this section we present further information about the precision experiments presented in Section IV-D. In Fig. 7 we show images of initial states from our experiments. During the 30 evaluations we did not move the target and always manually returned the gear to a similar position within several centimeters. In Table VII we present the means and spreads of points along the x and y directions from our experiments. These are the same results as in Fig. 6 to present an alternative quantitative analysis.

Appendix E Extracting motion from demonstrations

In section Section III-A we described how we split the motion into sections and how we extract active points. Since this is a part where we did task specific parameter selection we would like to describe the process in greater detail here. Assuming we already have clustering, this process has 2 main stages: 1) temporal segmentation and alignment or demonstrations 2) active point extraction.

Temporal segmentation. The core idea is that we noticed that for all the tasks we considered we can segment the motion based on the gripper actuation or external forces. We use these signals to construct a motion plan which is then executed by the controllers. To extract gripper actuation events we consider the gripper openness positions and note points where the position crosses a selected threshold. These time points are beginnings or ends of grasps. When manipulating small objects it was helpful to not open the gripper fully as that gave a higher grasp precision to the demonstrator and therefore we used a different threshold in each task. This threshold is noted in Table E as Gripper open. The second type of an event we consider is a beginning or an end of a force phase. To extract these we look at the vertical force measured by our force torque sensor and first smoothen the signal with a kernel with variance of 2.5 seconds. Then we define a local force maximum by computing flocal−maxf_{local-max} within a 5s windows. We then construct a normalized force signal as fnorm=f/max(fthreshold,flocal−max)f_{norm}=f/max(f_{threshold},f_{local-max}) where fthresholdf_{threshold} is noted in Table E as Max force. This ensures that our normalized forces are between 0 and 1, but the normalized forces stay small when no forces are being applied. We define a force event when the normalized force crosses a threshold of 0.5 as these will be points where we either engage in force-feedback behaviour or disengage from it.

When executing a force feedback behaviour we modify the visual servoing controller. Instead of following the vertical action from the visual servoing we apply a P controller on vertical component to keep a force of 1.5N1.5N. We also disallow advancing to the next frame unless the current force is within ±0.3N\pm 0.3N of this target.

In order to compare information across demonstrations we resample each segment of each demonstration into a fixed lenght block. However resampling linearly across time leads to oversampling of static phases. Therefore we resample the trajectories along the distance traveled. This is done using a per-timestep motion metric tm=∣∣vend−effector∣∣2/0.04+∣∣vfingers∣∣2/0.5t_{m}=||v_{end-effector}||^{2}/0.04+||v_{fingers}||^{2}/0.5 where vend−effectorv_{end-effector} is the linear velocity of the end-effector in m/sm/s and vfingersv_{fingers} captures the opening and closing of the gripper. tmt_{m} is then used to linearly interpolation within each motion segment to align all of the demonstrations.

The final step is to extract constant motion primitives which are needed because, after a grasp or a release of an object the gripper occludes large potion of the workspace. To do this we look at the standard deviation across the demonstrations at every time point vstdtv^{t}_{std}, the minimum vertical velocity vmintv^{t}_{min} and the maximum velocity reached during the demonstration vmaxv_{max}. If at any timepoint all demonstrations reliably contain the motion (i.e. vmint>0.5vstdtv^{t}_{min}>0.5v^{t}_{std}) and the motion is significant (i.e. vmint>0.03vmaxv^{t}_{min}>0.03v_{max}) we consider this a reliable signal that all demonstration contain the same motion. Active point extraction. As discussed in Section III-A we extract active points in several steps: We do a first pass in selecting active point candidates, cluster points and use points to vote for a clusters, and clean up the points within the selected clusters.

To do a first pass in selecting active points we use 3 seemingly trivial, but powerful heuristics: a) Reject points which are not currently visible b) Reject points which do not move c) Reject points which end up at different image locations. For a) we look at the final frame of the motion and define a threshold called Saliency in Table E for the fraction of demonstrations where the point must be visible in. Then for b) we reject points whose overall motion during the segment is less than a Is gripper fraction of the 90th quantile across the points. For c) we compute the variance of the final point location across demonstrations and require that it is less than a threshold. We called the threshold for the square root of the variance the Cross-demo variance.

After this procedure we have a set of points which belong the important object for a given segment. In many cases these points would be good enough to use as active points for the final visual servoing. One of the limitations of TAPIR, as most point tracking models, however is its ability to track points across very large changes of scale which is something we observe in our tasks as we use the gripper camera. Points chosen above are descriptive of the final frame which is often up-close and therefore those features don’t tend to be detected when the gripper is further out or when it is looking at the object from a different direction. To get points representative of the whole motion we leverage the clustering. We first clean the clusters by removing points which could not be assigned to any cluster with an error of less than Average Error removal. Then we count how many of the points belong to which cluster and normalize the counts by dividing by the largest number of votes. Then we merge all clusters which have score over a given threshold called Multi cluster. At this point we have a bank of relevant points which will work across all scales.

Lastly this clustering leads to a potentially large number of points. To remove them we compute an average visibility score across the whole motion and remove all points below this score. This threshold is called 2nd pass saliency in Table E. This last filtering step removes outliers leftover from the clustering process. To remove these we for each point computed the distance to the closest K point and averaged this across the frames where it was visible. Then we removed points where this was above a given threshold Dist. If at this point we had more than 128 points, we randomly selected which ones are going to be used during execution of the motion.

We believe that many of the hyper-parameters in Table E could be easily automatically extracted without the need to be manually specified by a human. However, in our case it was very easy to select and tune these by looking at visualisations from the intermediate steps of the process. For most of them we selected the value by looking at the distribution of the variable it was thresholding. Therefore we have not attempted to automate this process further at this stage.

Appendix F Controller Implementation Details

Given a set of target points gtg_{t} and corresponding detected points ptp_{t}, we aim to compute an action that would minimize the error, using a linear approximation of the function mapping actions to changes in ptp_{t}. We then take a step in that direction and repeat. This process can be summarized as:

However, in practice this approximation has two main issues: the optimal set of detected and target points to use for servoing may change depending on the current state, and the Jacobian of the error may not be a good approximation of the desired motion. These could likely be rectified by learning, but for simplicity, in this paper we use simple heuristics to resolve them analytically in ways that work for (approximately) rigid objects.

Point selection: TAPIR provides a visibility score visvis between 0 and 1. We use this score by summing the visibility across the demonstration frame which is being followed and the current frame and only using the 30% most visible or confidently-detected points.

Target selection: Recall that motion plans consist of a series of temporal segments of the demonstrations, along with a set of active points for each segment. Our low-level motion planner requires a single servoing target for every point at every timestep, but we have a full set of trajectories for each demonstration. How can we summarize this to a single point? In order to have precision, we would like to average across trajectories to reduce noise from the point tracking. However, simple averaging across the whole trajectory may not be sensible, since the initial configurations may be diverse. For many tasks (e.g. gluing), we must follow the entire trajectory. Thus, we take a hybrid approach: at the beginning of each temporal segment, we choose a single demonstration (the nearest one) and follow the full trajectory until the final frame of the demos, at which point we average across all available demos. This is a reliable procedure because we assume that our placement tasks form a ‘funnel’–i.e., the demonstrations may start off with objects in diverse states, but the object of interest always converges to the same location.

Specifically, we split every visual servoing motion into 2 phases: 1) demonstration follow phase and 2) the final phase. Phase (1) allows us to repeat the broad motion while phase (2) allows us to reach a consistent final state. At the beginning of every visual servoing phase we pick a demonstration to follow: we measure the Euclidean distance between the visible points on initial frame of every demonstration and the currently-detected locations, and pick the one with the lowest distance. We servo toward the initial location until the 30th percentile of per-point errors falls below a threshold, at which point we advance to the next frame of the same demonstration. We alternately servo when the error is above the threshold and advance when the error is below the threshold, until we reach the final frame of the demonstration, at which point the final phase begins. For the final phase (2), we want to achieve a consistency of behaviour no matter which demonstration was followed. We use the (visibility-weighted) average of points across demonstration as the target gg. To move on to the next timestep in phase 1 we use a threshold of 12 pixels and phase 2, we use a threshold of 2 pixels, but we multiply this threshold by 1.01 for every timestep we spend in this phase, thus gradually increasing the termination threshold to ensure that the temporal segment is guaranteed to terminate even in the presence of noise. Computing the Jacobian Given a set of target points, we must next compute a Jacobian with respect to the detected points, which will allow us to compute an action which will minimize the error. However, when the distance between the target points and the detected points is large, the linear approximation implied by the Jacobian may be poor: in particular, the model may use rotation and z-motion to explain errors due to in-plane translation. For example, when the object is far from the image center, the rotation and translation are highly correlated in the motions they will induce in the target points. Therefore, we use Gram-Schmidt to orthogonalize the last 4 columns of the Jacobian relative to the first two. Recall that our Jacobian can be written as:

Where uu is the x-position of a detected point, and vv is the y-position. Thus, the first two columns correspond to xx and yy translation, the third is zz translation, fourth and fifth are xx and yy tilt, and sixth is rotation. xx and yy translation dominates the loss, and the other columns are highly correlated with the first two. Columns 4 and 5 describing tilt are only used in the simulated ablations described in Section A. We orthogonalize the last 4 columns of the Jacobian with respect to the first two, after evaluating it on our points using the Gram–Schmidt Orthogonalisation method and denote it Jp⊥J^{\perp}_{p}. This leads to the final controller equation for the visual servoing command vvsv_{vs}:

Statistical Bias and the Jacobian As described above, a key problem with the Jacobian is statistical bias in which noise in the (2D) detections increases the estimate of action components which decrease the scale of the points (e.g. -z for a wrist-camera). We therefore modify the Jacobian so that it performs variance scaling for z-translation. Specifically, We flip the role of pp and gg in the solver, i.e. we compute both the standard Jacobian update vvsv_{vs}, and also the inverse Jacobian update vˉvs\bar{v}_{vs} which treats the demonstration points gg as controllable, and moves them toward the current states. The controller update is then (vvs−vˉvs)/2(v_{vs}-\bar{v}_{vs})/2. It is straightforward to show that this does not modify the vxv_{x} and vyv_{y} component of the motion, as the motion and inverse in-plane rotation are identical. Therefore, in this section, we focus on showing that the update for the vzv_{z} component becomes variance matching. To show this, we focus on only the Jacobian with respect to the zz component (recall that xx and yy are orthogonalized out from this). Thus, we have:

The zz-component of the Jacobian JpzJ^{z}_{p} is simply [−up−vp]⊤[-u_{p}-v_{p}]^{\top}, the −x-x and −y-y components of pp. The solution to this equation is Jpz−1⋅(p−g)J^{z-1}_{p}\cdot(p-g), where Jpz−1J^{z-1}_{p} is the pseudinverse of JpzJ^{z}_{p}, which can be computed analytically as −p/∣∣p∣∣2-p/||p||^{2}. In the presence of noise, we have that p=(p^+ϵp)p=(\hat{p}+\epsilon_{p}), where p^=[uc,vc]⊤\hat{p}=[u_{c},v_{c}]^{\top} is the true point locations, and ϵp\epsilon_{p} is the noise with variance σp2\sigma_{p}^{2} introduced by the detector in the current detections, which we assume is zero mean and uncorrelated with any other variables; we can define an analogous decomposition for the target points. Substituting and taking the expectation with respect to the noise yields an update of the form:

Because the ϵ\epsilon values are uncorrelated with other terms, if we consider the expected value, this simplifies to:

From this equation we can see that even if we reached the goal state, i.e. p^=g^\hat{p}=\hat{g}, our controller would be outputting a vz>0v_{z}>0. Specifically it would be outputting an action of 1/(1+g2/σp2)1/(1+g^{2}/\sigma_{p}^{2}) coming from the zoom-out bias. Computing the analogous term from a Jacobian in the other direction and subtracting it yields:

We can see that this expression is zero when ∣∣g∣∣2=∣∣p∣∣2||g||^{2}=||p||^{2} which can be interpreted as matching variances.

Appendix G RoboTAP Dataset

Video sources The videos come from DeepMind Robotics real-world data collections from Manipulation Task Suite (MTS). To make the data source diverse, we randomly sample a subset from the publicly released Sketchy dataset , S3K dataset and NIST dataset . The Sketchy data are sensed and recorded with 3 cage cameras and 3 wrist cameras (wide angle and depth), focusing on 3 variable-shape rigid objects coloured red, green and blue (rgb dataset) and 3 deformable objects: a soft ball, a rope and a cloth (deformable dataset). The S3K data contains a collection of rigid and nonrigid objects: e.g. shoes, toys. The NIST data contains gears and board that is widely known as the NIST challenge in Robotic Grasping and Manipulation . Figure 9 visualize the examples of RoboTAP videos and the annotated point tracks.

Annotation interface Figure 8 illustrates the annotation interface. The point annotation interface presented to annotators is based on . The interface loads and visualizes video in the visualization panel. Buttons in the video play panel allow annotators to navigate frames during annotation. The information panel provides basic information, e.g., the current and total number of frames. The annotation buttons (NEW and SUBMIT) allow annotators to add new point tracks or submit the labels if finished. The label info panel shows each annotated point track and the associated ‘tag’ string. The cursor is cross shaped which allows annotators to localize more precisely. Annotation process For acquisition of high-quality and diverse point tracks dataset, we encourage the annotators to choose any point that can be tracked reliably. Each track begins with an ENTER point, continues with MOVE points, and finishes when the annotator sets it as an EXIT point. Annotators can restart the track after an occlusion by adding another ENTER point and continuing. An optical flow based track assist algorithm is also utilized to interpolate the point tracks in the intermediate frames, this helps improve the annotation stability, accuracy and productivity. More details are provided in .

Dataset statistics Figure 10 shows more detailed statistics of RoboTAP dataset based on histograms. There are mainly 3 different video resolutions: 222 x 292, 480 x 640 and 768 x 1024. Although the total number of videos is not large, we find the point tracks cover a quite broad range of interests. For example, the majority of contiguous point tracks spans between 20 frames to 150 frames, with some longer than 1000 frames (e.g. 40 seconds). Among 11592 point tracks, around 5000 point tracks are visible across the episode, with the rest being occluded in a relatively uniform distribution. The most extremely moving points can reach a maximum distance that covers the whole video (i.e. from left to right) while a majority moves in a quarter of the image size.