VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-Training

Yecheng Jason Ma, Shagun Sodhani, Dinesh Jayaraman, Osbert Bastani, Vikash Kumar, Amy Zhang

Introduction

A long-standing challenge in robot learning is to develop robots that can learn a diverse and expanding set of manipulation skills from sensory observations (e.g., vision). This hope of developing general-purpose robots demands scalable and generalizable representation learning and reward learning to provide effective task representation and specification for downstream policy learning. Inspired by pre-training successes in computer vision (CV) (He et al., 2020; 2022) and natural language processing (NLP) (Devlin et al., 2018; Radford et al., 2019; 2021), pre-training visual representations on out-of-domain natural and human data (Deng et al., 2009; Grauman et al., 2022) has emerged as an effective solution for acquiring a general visual representation for robotic manipulation (Shah & Kumar, 2021; Parisi et al., 2022; Nair et al., 2022; Xiao et al., 2022) This paradigm is favorable to the traditional approach of in-domain representation learning because it does not require any intensive task-specific data collection or representation fine-tuning, and a single fixed representation can be used for a variety of unseen robotic domains and tasks (Parisi et al., 2022).

A key unsolved problem to pre-training for robotic control is the challenge of reward specification. Unlike simulated environments, real-world robotics tasks do not come with privileged environment state information or a well-shaped reward function defined over this state space. Prior pre-trained representations for control demonstrate results in only visual reinforcement learning (RL) in simulation, assuming access to a well-shaped dense reward function (Shah & Kumar, 2021; Xiao et al., 2022), or visual imitation learning (IL) from demonstrations (Parisi et al., 2022; Nair et al., 2022). In either case, substantial engineering effort is required for learning each new task. Instead, a simple and general way of specifying real-world manipulation tasks is by providing a goal image (Andrychowicz et al., 2017; Pathak et al., 2018) that captures the desired visual changes to the environment. However, as we demonstrate in our experiments, existing pre-trained visual representations do not produce effective reward functions in the form of embedding distance to the goal image, despite their effectiveness as pure visual encoders. Given that these models are already some of the most powerful models derived from computer vision, it begs the pertinent question of whether a universal visual reward function learned entirely from out-of-domain data is even possible.

In this paper, we show that such a general reward model can indeed be derived from a pre-trained visual representation, and we acquire this representation by treating representation learning from diverse human-video data as a big offline goal-conditioned reinforcement learning problem. Our idea philosophically diverges from all prior works: instead of taking what worked the best for CV tasks and “hope for the best” in visual control, we propose a more principled approach of using reinforcement learning itself as a pre-training mechanism for reinforcement learning. Now, this formulation certainly seems impractical at first because human videos do not contain any action information for policy learning. Our key insight is that instead of solving the impossible primal problem of direct policy learning from out-of-domain, action-free videos, we can instead solve the Fenchel dual problem of goal-conditioned value function learning. This dual value function, as we will show, can be trained without actions in an entirely self-supervised manner, making it suitable for pre-training on (out-of-domain) videos without robot action labels.

Theoretically, we show that this dual objective amounts to a novel form of implicit time contrastive learning, which attracts the representations of the initial and goal frame in the same trajectory, while implicitly repelling the representations of intermediate frames via recursive one-step temporal-difference minimization. These properties enable the representation to capture long-range temporal dependencies over distant task frames and inject local temporal smoothness over neighboring frames, making for smooth embedding distances that we show are the key ingredient of an effective reward function. This contrastive lens, importantly, enables the value function to be implicitly defined as a similarity metric in the embedding space, resulting in our simple final algorithm, Value-Implicit Pre-training (VIP); see Fig. 1 for an overview.

Trained on the large-scale, in-the-wild Ego4D human video dataset (Grauman et al., 2022) using a simple sparse reward, VIP is able to capture a general notion of goal-directed task progress that makes for effective reward-specification for unseen robot tasks specified via goal images. On an extensive set of simulated and real-robot tasks, VIP’s visual reward significantly outperforms those of prior pre-trained representations on a diverse set of reward-based policy learning paradigms. Coupled with a standard trajectory optimizer (Williams et al., 2017), VIP can solve ≈30%\approx 30\% of the tasks without any task-specific hyperparameter or representation fine-tuning and is the only representation that enables non-trivial progress on the set of more difficult tasks. Given more optimization budget, VIP’s performance can further improve up to ≈45%\approx 45\%, whereas other representations do worse due to reward hacking. When serving as both the visual representation and reward function for visual online RL, VIP again significantly outperforms prior methods by a wide-margin and achieves 40% aggregate success rate. To the best of our knowledge, these are the first demonstrations of a successful perceptual reward function learned using entirely out-of-domain data. Finally, we demonstrate that VIP can enable real-world few-shot offline RL with as few as 20 trajectories on a diverse suite of real-robot manipulation tasks, demonstrating for the first time that offline RL is possible in this low-data regime and paving the way towards truly scalable and autonomous robot learning.

Related Work

We review relevant literature on (1) Out-of-Domain Representation Pre-Training for Control, (2) Perceptual Reward Learning from Human Videos, and (3) Goal-Conditioned RL as Representation Learning. Due to space constraint, the latter two are included in App. B.

Out-of-Domain Representation Pre-Training for Control. Bootstrapping visual control using frozen representations pre-trained on out-of-domain non-robot data is a nascent field that has seen fast progress over the past year. Shah & Kumar (2021) demonstrates that pre-trained ResNet (He et al., 2016) representation on ImageNet (Deng et al., 2009) serves as effective visual backbone for simulated dexterous manipulation RL tasks. Parisi et al. (2022) finds ResNet models trained with unsupervised objectives, such as momentum contrastive learning (MOCO) (He et al., 2020), to surpass supervised objectives (e.g, image classification) for both visual navigation and control tasks. Xiao et al. (2022) demonstrates that masked-autoencoder (He et al., 2022) trained on diverse video data (Goyal et al., 2017; Shan et al., 2020) can be an effective visual embedding for online RL. The closest work to ours is R3M Nair et al. (2022), which is also pre-trained on the Ego4D dataset and attempts to capture temporal information in the videos by using time-contrastive learning (Sermanet et al., 2018); whereas VIP is fully self-supervised, R3M additionally requires video textual descriptions to align its representation. These prior works primarily re-purpose existing objectives and models for visual control and do not address the reward specification challenge. In contrast, VIP is the first to propose a novel RL-based objective for out-of-domain pre-training and is capable of producing generalizable dense reward signals that enable several new visuomotor control strategies that have not been demonstrated in this setting before.

Problem Setting and Background

In this section, we describe our problem setting of out-of-domain pre-training and provide formalism for downstream representation evaluation. Additional background on goal-conditioned reinforcement learning and contrastive learning is included in App. A.

Representation Evaluation. Given a choice of representation ϕ\phi, every evaluation task can be instantiated as a Markov decision process M(ϕ):=(ϕ(O),A,R(ot,ot+1;ϕ,g),T,γ,g)\mathcal{M}(\phi):=(\phi(O),A,R(o_{t},o_{t+1};\phi,g),T,\gamma,g), in which the state space is the induced space of observation embeddings, and the task is specified via a (set of) goal image(s) gg. Specifically, for a given transition tuple (ot,ot+1)(o_{t},o_{t+1}), we define the reward to be the goal-embedding distance difference (Lee et al., 2021; Li et al., 2022):

where Sϕ\mathcal{S}_{\phi} is a distance function in the ϕ\phi-representation space; in this work, we set Sϕ(ot;g):=−∥ϕ(ot)−ϕ(g)∥2\mathcal{S}_{\phi}(o_{t};g):=-\left\lVert\phi(o_{t})-\phi(g)\right\rVert_{2}. This reward function can be interpreted as a raw embedding distance reward with a reward shaping (Ng et al., 1999) term that encourages making progress towards the goal. This preserves the optimal policy but enables more efficient and robust policy learning.

Value-Implicit Pre-Training

In this section, we demonstrate how a self-supervised value-function objective can be derived from computing the dual of an offline RL objective on passive human videos (Section 4.1). Then, we show how this objective amounts to a novel implicit formulation of temporal contrastive learning (Section 4.2), which helps inducing a temporally smooth embedding favorable for downstream visual reward specification. Finally, we leverage this contrastive interpretation to instantiate a simple implementation (<10 lines of PyTorch code) of our dual value objective that does not explicitly learn a value network (Section 4.3), culminating in our final algorithm, Value-Implicit Pre-training (VIP).

While human videos are out-of-domain data for robots, they are in-domain for learning a goal-conditioned policy πH\pi_{H} over human actions, aH∼πH(ϕ(o)∣ϕ(g))a^{H}\sim\pi^{H}(\phi(o)\mid\phi(g)), for some human action space AHA^{H}. Therefore, given that human videos naturally contain goal-directed behavior, one reasonable idea of utilizing offline human videos for representation learning is to solve an offline goal-conditioned RL problem over the space of human policies and then extract the learned visual representation. To this end, we consider the following KL-regularized offline RL objective (Nachum et al., 2019) for some to-be-specified reward r(o,g)r(o,g):

Under assumption of deterministic transition dynamics, the dual optimization problem of equation 2 is

where μ0(o;g)\mu_{0}(o;g) is the goal-conditioned initial observation distribution, and D(o,o′;g)D(o,o^{\prime};g) is the goal-conditioned distribution of two consecutive observations in dataset DD.

2 Analysis: Implicit Time Contrastive Learning

While equation 3 will learn some useful visual representation via temporal value function optimization, in this section, we show that it can be understood as a novel implicit temporal contrastive learning objective that acquires temporally smooth embedding distance over video sequences, underpinning VIP’s efficacy jointly as a visual representation and reward for downstream control.

We begin by simplifying the expression in equation 3 by first assuming that the optimal V∗V^{*} is found:

where we have also re-written the maximization problem as a minimization problem. Now, after few algebraic manipulation steps (see App. C for a derivation), if we think of V∗(ϕ(o);ϕ(g))V^{*}(\phi(o);\phi(g)) as a similarity metric in the embedding space, then we can massage equation 4 into an expression that resembles the InfoNCE (Oord et al., 2018) time contrastive learning (Sermanet et al., 2018) (see App. A.2 for a definition and additional background) objective:

In particular, p(g)p(g) can be thought of the distribution of “anchor” observations, μ0(s;g)\mu_{0}(s;g) the distribution of “positive” samples, and D(o,o′;g)D(o,o^{\prime};g) the distribution of “negative” samples. Counter-intuitively and in contrast to standard single-view time contrastive learning (TCN), in which the positive observations are temporally closer to the anchor observation than the negatives, equation 5 has the positives to be as temporally far away as possible, namely the initial frame in the the same video sequence, and the negatives to be middle frames sampled in between. This departure is accompanied by the equally intriguing deviation of the lack of explicit repulsion of the negatives from the anchor; instead, they are simply encouraged to minimize the (exponentiated) one-step temporal-difference error in the representation space (the denominator in equation 5); see Fig. 1. Now, since the value function encodes negative discounted temporal distance, due to the recursive nature of value temporal-difference (TD), in order for the one-step TD error to be globally minimized along a video sequence, observations that are temporally farther away from the goal will naturally be repelled farther away in the representation space compared to observations that are nearby in time; in App. C.3, we formalize this intuition and show that this repulsion always holds for optimal paths. Therefore, the repulsion of the negative observations is an implicit, emergent property from the optimization of equation 5, instead of an explicit constraint as in standard (time) contrastive learning.

Now, we dive into why this implicit time contrastive learning is desirable. First, the explicit attraction of the initial and goal frames enables capturing long-range semantic temporal dependency as two frames that meaningfully indicate the beginning and end of a task are made close in the embedding space. This closeness is also well-defined due to the one-step TD backup that makes every embedding distance recursively defined to be the discounted number of timesteps to the goal frame. Combined with the implicit yet structured repulsion of intermediate frames, this push-and-pull mechanism helps inducing a temporally smooth and consistent representation. In particular, as we pass a video sequence in the training set through the trained representation, the embedding should be structured such that two trends emerge: (1) neighboring frames are close-by in the embedding space, (2) their distances to the last (goal) frame smoothly decrease due to the recursively defined embedding distances. To validate this intuition, in Fig. 2, we provide a simple toy example comparing implicit vs. standard time contrastive learning when trained on in-domain, task-specific demonstrations; details are included in App. E.2. As shown, standard time contrastive learning only enforces a coarse notion of temporal consistency and learns a non-locally smooth representation that exhibits many local minima. In contrast, VIP learns a much better structured embedding that is indeed temporally consistent and locally smooth.

3 Algorithm: Value-Implicit Pre-Training (VIP)

Recall that V∗V^{*} is assumed to be known for the derivation in Section 4.2, but in practice, its analytical form is rarely known. Now, given that V∗V^{*} plays the role of a distance measure in our implicit time contrastive learning framework, a simple and practical way to approximate V∗V^{*} is to set it to be a choice of similarity metric, bypassing having to explicitly parameterize it as a neural network. In this work, we choose the common choice of the negative L2L_{2} distance used in prior work Sermanet et al. (2018); Nair et al. (2022): V∗(ϕ(o),ϕ(g)):=−∥ϕ(o)−ϕ(g)∥2V^{*}(\phi(o),\phi(g)):=-\left\lVert\phi(o)-\phi(g)\right\rVert_{2}. Given this choice, our final representation learning objective is as follows:

in which we also absorb the exponent of the log-sum-exp term in 4 into the inner exp⁡(⋅)\exp(\cdot) term via Jensen’s inequality; we found this upper bound to be numerically more stable. To sample video trajectories from DD, because any sub-trajectory of a video is also a valid video sequence, VIP samples these sub-trajectories and treats their initial and last frames as samples from the goal and initial-state distributions (Step 3 in Alg. 1). Altogether, VIP training is illustrated in Alg. 1; it is simple and its core training loop can be implemented in fewer than 10 lines of PyTorch code (Alg. 2 in App. D.3).

Experiments

In this section, we demonstrate VIP’s effectiveness as both a pre-trained visual reward and representation on three distinct reward-based policy learning settings. We begin by detailing VIP’s training and introducing baselines. Then, we present the full evaluation results for each setting, and we conclude with qualititave analysis, delving into VIP’s unique effectiveness.

VIP Training. We use a standard ResNet50 (He et al., 2016) architecture as VIP’s visual backbone and train on a subset of the Ego4D dataset (Grauman et al., 2022), a large-scale egocentric video dataset consisting of humans accomplishing diverse tasks around the world. These choices are identical to a prior work (Nair et al., 2022); additionally, we use the exact same hyperparameters (e.g., batch size, optimizer, learning rate) as in Nair et al. (2022). See App. D for details.

Baselines. The closest comparison is R3M (Nair et al., 2022), which pre-trains on the same Ego4D dataset using a combination of time contrastive learning, L1L_{1} weight regularization, and language embedding consistency losses. We also consider a self-supervised ResNet50 network trained on ImageNet using Momentum Contrastive (MoCo), a supervised ResNet50 network trained on ImageNet, and CLIP Radford et al. (2021), covering a wide range of pre-existing visual representations that have been used for robotics control (Shah & Kumar, 2021; Parisi et al., 2022; Cui et al., 2022), though none has been tested in our three reward-based settings in which the reward also has to be produced by the representation. Note that the visual backbone for all methods is ResNet50, enabling a fair comparison of the pre-training objectives. Besides VIP, all models are taken from their publicly released checkpoints. In Appendix G.1, we additionally compare to MoCo and Masked Auto-Encoder (MAE) models trained on Ego4D.

Evaluation Environments For our simulation experiments, we consider the FrankaKitchen (Gupta et al., 2019) environment, in which a 7-DoF Franka robot is tasked with manipulating common household kitchen objects to pre-specified configurations. We use all 12 subtasks supported in the environment and 3 camera views (left, center, right) for each task, in total of 36 visual manipulation tasks. Furthermore, we consider two initial robot state for every task, one Easy setting in which the end-effector is initialized close to the object of interest, and one Hard setting in which the end-effector is uniformly initialized above the stove edge regardless of the task. The task horizon is 50 (resp. 100) for the Easy (resp. Hard) setting. Each task is specified via a goal image; see Fig. 3 for example goal images for all three views, and see Fig. 8-9 in App. E.1 for initial and goal frames for all tasks).

We evaluate pre-trained representations’ capability as pure visual reward functions by using them to directly synthesize a sequence of actions using trajectory optimization In particular, we use model-predictive path-integral (MPPI) (Williams et al., 2017). To evaluate each proposed sequence of actions, we directly roll it out in the simulator for simplicity. Vanilla trajectory optimization, though sample efficient, is prone to local minima in the reward landscape due to a lack of trial-and-error exploration. We hypothesize that online RL may be able to overcome bad local minima, but it comes with the added challenge of demanding the pre-trained representation to provide both the visual reward and representation for learning a closed-loop policy. To further study the importance of a learned dense reward in online RL, we compare to using ground-truth task sparse reward coupled with VIP’s visual representation, VIP (Sparse). The RL algorithm we use is natural policy gradient (NPG) (Kakade, 2001). We leave all experiment details are in App. E.3-E.4. In Fig. 4, we report each representation’s cumulative success rate averaged over all configurations (3 seeds * 3 cameras * 12 tasks = 108 runs); the success rate for online RL is computed over a separate set of test rollouts.

Examining the MPPI results, we see that VIP is substantially better than all baselines in both Easy and Hard settings, and is the only representation that makes non-trivial progress on the Hard setting. In Fig. 5, we couple VIP and the strongest baselines (R3M, Resnet) with increasingly more powerful MPPI optimizers (i.e., more trajectories per optimization step; default 3232 are used in Fig. 4). As shown, while VIP steadily benefits from stronger optimizers and can reach an average success rate of 44%, baselines often do worse when MPPI is given more compute budget, suggesting that their reward landscapes are filled with local minima that do not correlate with task progress and are easily exploited by stronger optimizers. To further validate these observations, in App. E.3, we report the L2L_{2} error to the ground-truth goal-image robot and object poses after each environment step taken by MPPI. We find that VIP is able to minimize both the robot and object pose errors quite robustly over diverse tasks and views, whereas several baselines in fact increase the robot pose error on average. Given these findings, we hypothesize that VIP’s reward functions are able capture task-salient information in the visual observations. In App. G.4, we validate this hypothesis and find that on at least one camera view for 8 out of the 12 tasks, VIP’s rewards are highly correlated with the human-engineered state-based dense rewards, with correlation coefficients as high as R2=0.95R^{2}=\textbf{0.95}, highlighting its potential of replacing manual reward engineering without any prior knowledge about the robot domain or tasks.

Switching gears to online RL, VIP again achieves consistently superior performance. VIP (Sparse)’s inability to solve any task, despite a strong visual representation provided by VIP itself, indicates the necessity of dense reward in solving these challenging visual manipulation tasks and further accentuates VIP’s versatility doubling as both visual reward and representation; furthermore, we find that sparse reward coupled with even the true state representation is unable to make any progress. Finally, we comment that whereas sparse reward still requires human engineering via installing additional sensors (Rajeswar et al., 2021; Singh et al., 2019) and faces exploration challenges (Nair et al., 2018), with VIP, the end-user has to provide only a goal image.

2 Real-World Few-Shot Offline Reinforcement Learning

In this section, we demonstrate how VIP’s reward and representation can power a simple and practical system for real-world robot learning in the form of few-shot offline reinforcement learning, making offline RL simple, sample-efficient, and more effective than BC with almost no added complexity.

3 Qualitative Analysis

We hypothesize that VIP learns the most temporally smooth embedding that enables effective zero-shot reward-specification, and present several qualitative experiments investigating this claim. First, in Fig. 6, we visually overlay and compare the embedding distance-to-goal curves for each representation on representative videos from both Ego4D and our real-robot dataset; the curve for every representation is normalized to have initial distance 11 to enable comparison. As shown, VIP has the most visually smooth curves, whereas all other methods exhibit “bumps” (i.e., positive slope at a step) that signal more prevalent presence of local minima in their reward landscapes; in App. G.5, we provide additional embedding curves. Finally, we quantify the total number of “bumps” each representation encounters over both datasets in App. G.6, and find VIP indeed has much fewer bumps.

Conclusion

We have proposed Value-Implicit Pre-training (VIP), a self-supervised value-based pre-training objective that is highly effective in providing both the visual reward and representation for downstream unseen robotics tasks. VIP is derived from first principles of dual reinforcement learning and admits an appealing connection to an implicit and more powerful formulation of time contrastive learning, which captures long-range temporal dependency and injects local temporal smoothness in the representation to make for effective zero-shot reward specification. Trained entirely on diverse, in-the-wild human videos, VIP demonstrates significant gains over prior state-of-art pre-trained visual representations on an extensive set of policy learning settings. Notably, VIP can enable sample-efficient real-world offline RL with just handful of trajectories. Altogether, we believe that VIP makes an important contribution in both the algorithmic frontier of visual pre-training for RL and practical real-world robot learning.

Acknowledgment

This work was done while YJM was an intern at Meta AI. We thank Aravind Rajeswaran and members of Meta AI for helpful discussions, Jay Vakil for assistance on the real-robot experiments, and Vincent Moens for assistance on releasing the model.

Author Contributions

JYM developed VIP, designed and implemented models and experiments, led the real-robot experiments, and wrote the paper. SS oversaw the project logistics, assisted on the software infrastructure and simulation experiments, and edited the writing. DJ and OB provided feedback on the project. VK oversaw and advised on the project, assisted on the software infrastructure, designed experiments, and supervised the real-robot experiments. AZ oversaw and advised on the project, designed experiments, and guided and edited the paper writing.

Reproducibility Statement

We have open-sourced code for using our pre-trained VIP model and training a new VIP model using any custom video dataset at https://github.com/facebookresearch/vip; the instruction for model training and inference is included in the README.md file in the supplementary file, and the hyperparameters are already configured. A PyTorch-based pseudocode for VIP is also included in Alg. 2.

References

Part I Appendix

This section is adapted from Ma et al. (2022b). We consider a goal-conditioned Markov decision process from visual state space: M=(O,A,G,r,T,μ0,γ)\mathcal{M}=(O,A,G,r,T,\mu_{0},\gamma) with state space OO, action space AA, reward r(o,g)r(o,g), transition function o′∼T(o,a)o^{\prime}\sim T(o,a), the goal distribution p(g)p(g), and the goal-conditioned initial state distribution μ0(o;g)\mu_{0}(o;g), and discount factor γ∈(0,1]\gamma\in(0,1]. We assume the state space OO and the goal space GG to be defined over RGB images. The objective of goal-conditioned RL is to find a goal-conditioned policy π:O×G→Δ(A)\pi:O\times G\rightarrow\Delta(A) that maximizes the discounted cumulative return:

The goal-conditioned state-action occupancy distribution dπ(o,a;g):O×A×G→d^{\pi}(o,a;g):O\times A\times G\rightarrow of π\pi is

which captures the goal-conditioned visitation frequency of state-action pairs for policy π\pi. The state-occupancy distribution then marginalizes over actions: dπ(o;g)=∑adπ(o,a;g)d^{\pi}(o;g)=\sum_{a}d^{\pi}(o,a;g). Then, it follows that π(a∣o,g)=dπ(o,a;g)dπ(o;g)\pi(a\mid o,g)=\frac{d^{\pi}(o,a;g)}{d^{\pi}(o;g)}. A state-action occupancy distribution must satisfy the Bellman flow constraint in order for it to be an occupancy distribution for some stationary policy π\pi:

A.2 InfoNCE & Time Contrastive Learning.

As VIP can be understood as a implicit and smooth time contrastive learning objective, we provide additional background on the InfoNCE Oord et al. (2018) and time contrastive learning (TCN) (Sermanet et al., 2018) objective to aid comparison in Section 4.2.

TCN is a contrastive learning objective that learns a representation that in timeseries data (e.g., video trajectories). The original work (Sermanet et al., 2018) considers multi-view videos and perform contrastive learning over frames in separate videos; in this work, we consider the single-view variant. At a high level, TCN attracts representations of frames that are temporally close, while pushing apart those of frames that are farther apart in time. More precisely, given three frames sampled from a video sequence (ot1,ot2,ot3)(o_{t_{1}},o_{t_{2}},o_{t_{3}}), where t1<t2<t3t_{1}<t_{2}<t_{3}, TCN would attract the representations of ot1o_{t_{1}} and ot2o_{t_{2}} and repel the representation of ot3o_{t_{3}} from ot1o_{t_{1}}. This idea can be formally expressed via the following objective:

Given a “positive” window of KK steps and a uniform distribution among valid positive samples, we can write equation 12 as

in which each term inside the expectation is a standalone InfoNCE objective tailored to observation sequence data.

Appendix B Extended Related Work

Perceptual Reward Learning from Human Videos. Human videos provide a rich natural source of reward and representation learning for robotic learning. Most prior works exploit the idea of learning an invariant representation between human and robot domains to transfer the demonstrated skills (Sermanet et al., 2016; 2018; Schmeckpeper et al., 2020; Chen et al., 2021; Xiong et al., 2021; Zakka et al., 2022; Bahl et al., 2022). However, training these representations require task-specific human demonstration videos paired with robot videos solving the same task, and cannot leverage the large amount of “in-the-wild” human videos readily available. As such, these methods require robot data for training, and learn rewards that are task-specific and do not generalize beyond the tasks they are trained on. In contrast, VIP do not make any assumption on the quality or the task-specificity of human videos and instead pre-trains an (implicit) value function that aims to capture task-agnostic goal-oriented progress, which can generalize to completely unseen robot domains and tasks.

Goal-Conditioned RL as Representation Learning. Our pre-training method is also related to the idea of treating goal-conditioned RL as representation learning. Chebotar et al. (2021) shows that a goal-conditioned Q-function trained with offline in-domain multi-task robot data learns an useful visual representation that can accelerate learning for a new downstream task in the same domain. Eysenbach et al. (2022) shows that goal-conditioned Q-learning with a particular choice of reward function can be understood as performing contrastive learning. In contrast, our theory introduces a new implicit time contrastive learning, and states that for any choice of reward function, the dual formulation of a regularized offline GCRL objective can be cast as implicit time contrast. This conceptual bridge also explains why VIP’s learned embedding distance is temporally smooth and can be used as an universal reward mechanism. Finally, whereas these two works are limited to training on in-domain data with robot action labels, VIP is able to leverage diverse out-of-domain human data for visual representation pre-training, overcoming the inherent limitation of robot data scarcity for in-domain training.

Our work is also closely related to Ma et al. (2022b), which first introduced the dual offline GCRL objective based on Fenchel duality (Rockafellar, 1970; Nachum & Dai, 2020; Ma et al., 2022a). Whereas Ma et al. (2022b) assumes access to the true state information and focuses on the offline GCRL setting using in-domain offline data with robot action labels, we extend the dual objective to enable out-of-domain, action-free pre-training from human videos. Our particular dual objective also admits a novel implicit time contrastive learning interpretation, which simplifies VIP’s practical implementation by letting the value function be implicitly defined instead of a deep neural network as in Ma et al. (2022b).

Appendix C Technical Derivations and Proofs

We first reproduce Proposition 4.1 for ease of reference:

Under assumption of deterministic transition dynamics, the dual optimization problem of

where μ0(o;g)\mu_{0}(o;g) is the goal-conditioned initial observation distribution, and D(o,o′;g)D(o,o^{\prime};g) is the goal-conditioned distribution of two consecutive observations in dataset DD.

Fixing a choice of ϕ\phi, the inner optimization problem operates over a ϕ\phi-induced state and goal space, giving us equation 16. Then, applying Proposition 4.2 of Ma et al. (2022b) to the inner optimization problem, we immediately obtain

Finally, sampling embedded states from dD(ϕ(o),ϕ(o′);ϕ(g))d^{D}(\phi(o),\phi(o^{\prime});\phi(g)) is equivalent to sampling from D(o,o′;g)D(o,o^{\prime};g), assuming there is no embedding collision (i.e., ϕ(o)≠ϕ(o′),∀o≠o′\phi(o)\neq\phi(o^{\prime}),\forall o\neq o^{\prime}), which can be satisfied by simply augmenting any ϕ\phi by concatenating the input to the end. Then, we have our desired expression:

C.2 VIP Implicit Time Contrast Learning Derivation

This section provides all intermediate steps to go from equation 4 to equation 5. First, we have

We can equivalently write this objective as

C.3 VIP Implicit Repulsion

In this section, we formalize the implicit repulsion property of VIP objective (equation 5); in particular, we prove that under certain assumptions, it always holds for optimal paths.

Suppose V∗(s;g):=−∥ϕ(s)−ϕ(g)∥2V^{*}(s;g):=-\left\lVert\phi(s)-\phi(g)\right\rVert_{2} for some ϕ\phi, under the assumption of deterministic dynamics (as in Proposition 4.1), for any pair of consecutive states reached by the optimal policy, (st,st+1)∼π∗(s_{t},s_{t+1})\sim\pi^{*}, we have that

A proof can be found in Section 1.1.3 of Agarwal et al. (2019). Then, due to the Bellman optimality equation, we have that

Given that the dynamics is deterministic and equation 24, we have that

Now, for (st,at,st+1)∼π∗(s_{t},a_{t},s_{t+1})\sim\pi^{*}, this further simplifies to

The implication of this result is that at least along the trajectories generated by the optimal policy, the representation will have monotonically decreasing and well-behaved embedding distances to the goal. Now, since in practice, VIP is trained on goal-directed (human video) trajectories, which are near-optimal for goal-reaching, we expect this smoothness result to be informative about VIP’s embedding practical behavior and help formalize out intuition about the mechanism of implicit time contrastive learning. As confirmed by our qualitative study in Section 5.3, We highlight that VIP’s embedding is indeed much smoother than other baselines along test trajectories on both Ego4D and on our real-robot dataset. This smoothness along optimal paths makes it easier for the downstream control optimizer to discover these paths, conferring VIP representation effective zero-shot reward-specification capability that is not attained by any other comparison.

Appendix D VIP Training Details

D.2 VIP Hyperparameters

Hyperparameters used can be found in Table 2.

D.3 VIP Pytorch Pseudocode

In this section, we present a pseudocode of VIP written in PyTorch (Paszke et al., 2019), Algorithm 2. As shown, the main training loop can be as short as 10 lines of code.

Appendix E Simulation Experiment Details.

In this section, we describe the FrankaKitchen suite for our simulation experiments. We use 12 tasks from the v0.1 versionhttps://github.com/vikashplus/mj_envs/tree/v0.1real/mj_envs/envs/relay_kitchen of the environment.

We use the environment default initial state as the initial state and frame for all tasks in the Hard setting. In the Easy setting, we use the 20th frame of a demonstration trajectory and its corresponding environment state as the initial frame and state. The goal frame for both settings is chosen to be the last frame of the same demonstration trajectory. The initial frames and goal frame for all 12 tasks and 3 camera views are illustrated in Figure 8-9. In the Easy setting, the horizon for all tasks is 50 steps; in the Hard setting, the horizon is 100 steps. Note that using the 20th frame as the initial state is a crude way for initializing the robot, and for some tasks, this initialization makes the task substantially easier, whereas for others, the task is still considerably difficult. Furthermore, some tasks become naturally more difficult depending on camera viewpoints. For these reasons, it is worth noting that our experiment’s emphasis is on the aggregate behavior of pre-trained representations, instead of trying to solve any particular task as well as possible.

E.2 In-Domain Representation Probing

In this section, we describe the experiment we performed to generate the in-domain VIP vs. TCN comparison in Figure 2. We fit VIP and TCN representations using 100 demonstrations from the FrankaKitchen sdoor_open task (center view). For TCN, we use R3M’s implementation of the TCN loss without any modification; this also allows our findings in Figure 2 to extend to the main experiment section. The visual architecture is ResNet34, and the output dimension is 22, which enables us to directly visualize the learned embedding. Different from the out-of-domain version of VIP, we also do not perform weight penalty, trajectory-level random cropping data augmentation, or additional negative sampling. Besides these choices, we use the same hyperparameters as in Table 2 and train for 2000 batches.

E.3 Trajectory Optimization

We use a publicly available implementation of MPPIhttps://github.com/aravindr93/trajopt/blob/master/trajopt/algos/mppi.py, and make no modification to the algorithm or the default hyperparameters. In particular, the planning horizon is 1212 and 32 sequences of actions are proposed per action step. Because the embedding reward (equation 1) is the goal-embedding distance difference, the score (i.e., sum of per-transition reward) of a proposed sequence of actions is equivalent to the negative embedding distance (i.e., Sϕ(ϕ(oT);ϕ(g))S_{\phi}(\phi(o_{T});\phi(g))) at the last observation.

In this section, we visualize the per-step robot and object pose L2L_{2} error with respect to the goal-image poses. We report the non-cumulative curves (on the success rate as well) for more informative analysis.

E.4 Reinforcement Learning

We use a publicly available implementation of NPGhttps://github.com/aravindr93/mjrl/blob/master/mjrl/algos/npg_cg.py, and make no modification to the algorithm or the default hyperparameters. In the Easy (resp. Hard) setting, we train the policy until 500000 (resp. 1M) real environment steps are taken. For evaluation, we report the cumulative maximum success rate on 50 test rollouts from each task configuration (50*108=5400 total rollouts) every 10000 step.

Appendix F Real-World Robot Experiment Details

The robot learning environment is illustrated in Figure 11; a RealSense camera is mounted on the right edge of the table, and we only use the RGB image stream without depth information for data collection and policy learning.

Each task is specified via a set of goal images that are chosen to be the last frame of all demonstrations for the task. Hence, the goal embedding used to compute the embedding reward (equation 1( for each task is the average over the embeddings of all goal frames.

The tasks (in their initial positions) using a separate high-resolution phone camera are visualized in Figure 12. Sample demonstrations in the robot camera view are visualized in Figure 13.

F.2 Training and Evaluation Details

The policy network is implemented as a 2-layer MLP with hidden sizes $$. As in R3M’s real-world robot experiment setup, the policy takes in concatenated visual embedding of current observation and robot’s proprioceptive state and outputs robot action. The policy is trained with a learning rate of 0.001, and a batch size of 32 for 20000 steps.

For RWR’s temperature scale, we use τ=0.1\tau=0.1 for all tasks, except CloseDrawer where we find τ=1\tau=1 more effective for both VIP and R3M.

For policy evaluation, we use 10 test rollouts with objects randomly initialized to reflect the object distribution in the expert demonstrations. The rollout horizon is 100 steps.

F.3 Additional Analysis & Context

Offline RL vs. imitation learning for real-world robot learning. Offline RL, though known as the data-driven paradigm of RL (Levine et al., 2020), is not necessarily data efficient (Agarwal et al., 2021), requiring hundreds of thousands of samples even in low-dimensional simulated tasks, and requires a dense reward to operate most effectively (Mandlekar et al., 2021; Yu et al., 2022). Furthermore, offline RL algorithms are significantly more difficult to implement and tune compared to BC (Kumar et al., 2021; Zhang & Jiang, 2021). As such, the dominant paradigm of real-world robot learning is still learning from demonstrations (Jang et al., 2022; Mandlekar et al., 2018; Ebert et al., 2021). With the advent of VIP-RWR, offline RL may finally be a practical approach for real-world robot learning at scale.

Performance of R3M-BC. Our R3M-BC, though able to solve some of the simpler tasks, appears to perform relatively worse than the original R3M-BC in Nair et al. (2022) on their real-world tasks. To account for this discrepancy, we note that our real-world experiment uses different software-hardware stacks and tasks from the original R3M real-world experiments, so the results are not directly comparable. For instance, camera placement, an important variable for real-world robot learning, is chosen differently in our experiment and that of R3M; in R3M, a different camera angle is selected for each task, whereas in our setup, the same camera view is used for all tasks. Furthermore, we emphasize that our focus is not the absolute performance of R3M-BC, but rather the relative improvement R3M-RWR provides on top of R3M-BC.

F.4 Qualitative Analysis

In this section, we study several interesting policy behaviors VIP-RWR acquire. Policy videos are included in our supplementary video.

Robust key action execution. VIP-RWR is able to execute key actions more robustly than the baselines; this suggests that its reward information helps it identify necessary actions. For example, as shown in Figure 14, on the PickPlaceMelon task, failed VIP-RWR rollouts at least have the gripper grasp onto the watermelon, whereas for other baselines, the failed rollouts do not have the watermelon between the gripper and often incorrectly push the watermelon to touch the plate’s outer edge, preventing pick-and-place behavior from being executed.

Task re-attempt. We observe that VIP-RWR often learns more robust policies that are able to perform recovery actions when the task is not solved on the first attempt. For instance, in both CloseDrawer and FoldTowel, there are trials where VIP-RWR fails to close the drawer all the way or pick up the towel edge right away; in either case, VIP-RWR is able to re-attempt and solves the task (see our supplementary video). This is a known advantage of offline RL over BC (Kumar et al., 2022; Levine et al., 2020); however, we only observe this behavior in VIP-RWR and not R3M-RWR, indicating that this advantage of offline RL is only realized when the reward information is sufficiently informative.

Appendix G Additional Results

In this section, we compare to two additional representations trained on the same pre-training dataset of Ego4D. We compare them to VIP and R3M in the trajectory optimization setting and evaluate their performance as the optimization budget increases in the spirit of Figure 5. The result are shown in Figure 15. Consistent with the findings in the main text, the pre-training dataset is not the source of VIP’s empirical gains; all prior state-of-art pre-training methods struggle as zero-shot reward functions. Notably, MAE uses a vision transformer (Dosovitskiy et al., 2020) architecture backbone and exhibits an improving trend in performance similar to VIP; however, MAE’s absolute performance is still far inferior to VIP.

G.2 Value-Based Pre-Training Ablation: Least-Square Temporal-Difference

While VIP is the first value-based pre-training approach and significantly outperforms all existing methods, we show that this effectiveness is also unique to VIP and not to training a value function. To this end, we show that a simpler value-based baseline does not perform as well. In particular, we consider Least-Square Temporal-Difference policy evaluation (LSTD) (Bradtke & Barto, 1996; Sutton & Barto, 2018) to assess the importance of the choice of value-training objective:

in which we also parameterize VV as the negative L2L_{2} embedding distance as in VIP. Given that human videos are reasonably goal-directed, the value of the human behavioral policy computed via LSTD should be a decent choice of reward; however, LSTD does not capture the long-range dependency of initial to goal frames (first term in equation 3), nor can it obtain a value function that outperforms that of the behavioral policy. We train LSTD using the exact same setup as in VIP, differing in only the training objective, and compare it against VIP in our trajectory optimization settings.

As shown in Fig. 16, interestingly, LSTD already works better than all prior baselines in the Easy setting, indicating that value-based pre-training is indeed favorable for reward-specification. However, its inability to capture long range temporal dependency as in VIP (the first term in VIP’s objective) makes it far less effective on the Hard setting, which require extended smoothness in the reward landscape to solve given the distance between the initial observation and the goal. These results show that VIP’s superior reward specification comes precisely from its ability to capture both long-range temporal dependencies and local temporal smoothness, two innate properties of its dual value objective and the associated implicit time contrastive learning interpretation. To corroborate these findings, we have also included LSTD in our qualitative reward curve and histogram analysis in App. G.5, G.7, and G.8 and finds that VIP generates much smoother embedding than LSTD.

G.3 Visual Imitation Learning

One alternative hypothesis to VIP’s smoother embedding for its superior reward-specification capability is that it learns a better visual representation, which then naturally enables a better visual reward function. To investigate this hypothesis, we compare representations’ capability as a pure visual encoder in a visual imitation learning setup. We follow the training and evaluation protocol of (Nair et al., 2022) and consider 12 tasks combined from FrankaKitchen, MetaWorld (Yu et al., 2020), and Adroit (Rajeswaran et al., 2017), 3 camera views for each task, and 3 demonstration dataset sizes, and report the aggregate average maximum success rate achieved during training. R3M-Lang is the publicly released R3M variant without supervised language training. The average success rates over all tasks are shown in Table 4; the letter inside ()() stands for the pre-training dataset with EE referring to Ego4D and II Imagenet.

These results suggest that with current pre-training methods, the performance on visual imitation learning may largely be a function of the pre-training dataset, as all methods trained on Ego4D, even our simple baseline LSTD, performs comparably and are much better than the next best baseline not trained on Ego4D. Conversely, this result also suggests that despite not being designed for this purely supervised learning setting, value-based approaches constitute a strong baseline, and VIP is in fact currently the state-of-art for self-supervised methods. While these results highlight that VIP is effective even as a pure visual encoder, a necessary requirement for joint effectiveness for visual reward and representation, it fails to explain why VIP is far superior to R3M in reward-based policy learning. As such, we conclude that studying representations’ capability as a pure visual encoder may not be sufficient for distinguishing representations that can additionally perform zero-shot reward-specification.

G.4 Embedding and True Rewards Correlation

In this section, we create scatterplots of embedding reward vs. true reward on the trajectories MPPI have generated to assess whether the embedding reward is correlated with the ground-truth dense reward. More specifically, for each transition in the MPPI trajectories in Figure 4, we plot its reward under the representation that was used to compute the reward for MPPI versus the true human-crafted reward computed using ground-truth state information. The dense reward in FrankaKitchen tasks is a weighted sum of (1) the negative object pose error, (2) the negative robot pose error, (3) bonus for robot approaching the object, and (4) bonus for object pose error being small. This dense reward is highly tuned and captures human intuition for how these tasks ought to be best solved. As such, high correlation indicates that the embedding is able to capture both intuitive robot-centric and object-centric task progress from visual observations. We only compare VIP and R3M here as a proxy for comparing our implicit time contrastive mechanism to the standard time contrastive learning.

The scatterplots over all tasks and camera views (Easy setting) are shown in Figure 17,18, and 19. VIP rewards exhibit much greater correlation with the ground-truth reward on its trajectories that do accomplish task, indicating that when VIP does solve a task, it is solving the task in a way that matches human intuition. This is made possible via large-scale value pre-training on diverse human videos, which enables VIP to extract a human notion of task-progress that transfers to robot tasks and domains. These results also suggest that VIP has the potential of replacing manual reward engineering, providing a data-driven solution to the grand challenge of reward engineering for manipulation tasks. However, VIP is not yet perfect in its current form. Both methods exhibit local minima where high embedding distances in fact map to lower true rewards; however, this phenomenon is much severe for R3M. On 8 out of 12 tasks, VIP at least has one camera view in which its rewards are highly correlated with the ground-truth rewards on its MPPI trajectories.

G.5 Embedding Distance Curves

In Figure 20, we present additional embedding distance curves for all methods on Ego4D and our real-robot offline RL datasets. For Ego4D, we randomly sample 4 videos of 50-frame long (see Appendix G.6 for how these short snippets are sampled), and for our robot dataset, we compute the embedding distance curves for the 4 sample demonstrations in Figure 13. As shown, on all tasks in the real-robot dataset, VIP is distinctively more smooth than any other representation. This pattern is less accentuated on Ego4D. This is because a randomly sampled 50-frame snippet from Ego4D may not coherently represent a task solved from beginning to completion, so an embedding distance curve is not inherently supposed to be smoothly declining. Nevertheless, VIP still exhibits more local smoothness in the embedding distance curves, and for the snippets that do solve a task (the first two videos), it stands out as the smoothest representation.

G.6 Embedding Distance Curve Bumps

In this section, we compute the fraction of negative embedding rewards (equivalently, positive slopes in embedding embedding distance curves) for each video sequence and average over all video sequences in a dataset. Each sequence in our robot dataset is of 50 frames, and we use each sequence without any further truncation. For Ego4D, video sequences are of variable length. For each long sequence of more than 50 frames, we use the first 50 frames. We do not include videos shorter than 50 frames, in order to make the average fraction for each representation comparable between the two distinct datasets. Note that for Ego4D, due to its in-the-wild nature, it is not guaranteed that a 50-frame segment represents one task being solved from beginning to completion, so there may be naturally bumps in the embedding distance curve computed with respect to the last frame, as earlier frames may not actually be progressing towards the last frame in a goal-directed manner.The full results are shown in Table 5. VIP has fewest bumps in Ego4D videos, and this notion of smoothness transfer to the robot dataset. Furthermore, since the robot videos are in fact visually simpler and each video is guaranteed to be solving one task, the bump rate is actually lower despite the domain gap. While this observation generally also holds true for other representations, it notably does not hold for R3M, which is trained using standard time contrastive learning.

G.7 Embedding Reward Histograms (Real-Robot Dataset)

We present the reward histogram comparison against all baselines in Figure 21. The trend of VIP having more small, positive rewards and fewer extreme rewards in either direction is consistent across all comparisons.

G.8 Embedding Reward Histograms (Ego4D)

We present the reward histogram comparison against all baselines in Figure 22. The histograms are computed using the same set of 50-frame Ego4D video snippets as in Appendix G.6. The y-axis is in log-scale due to the large total count of Ego4D frames. As discussed, Ego4D video segments are less regular than those in our real-robot dataset, and this irregularity contributes to all representations having significantly more negative rewards compared to their histograms on the real-robot dataset. Nevertheless, the relative difference ratio’s pattern is consistent, showing VIP having far more rewards that lie in the first positive bin. Furthermore, VIP also has significantly fewer extreme negative rewards compared to all baselines.

Appendix H Limitations and Future Work

In this section, we describe limitations within the current VIP formulation and model and some potential future directions.

VIP is currently limited to providing rewards for tasks that can be specified via a goal image. While this encompasses a wide range of robotics tasks, many tasks cannot be fully expressed via a static image, such as ones that require following intermediate instructions and steps. Likewise, though not a strict assumption, we have only tested VIP with visual goals from the same domain (robots are not necessarily in the goal image). Extending VIP to be compatible with even more flexible and extensive forms of goals is a fruitful direction for expanding VIP’s capability.

VIP current parameterizes the value function as a symmetric embedding distance; this assumes that the environment is reversible (i.e., it is equally easy to get from oo to gg and from gg to oo), which may not hold in practice. While we did not observe this to affect practical performance, we may improve performance by parameterizing V(o;g)V(o;g) as some distance function that supports asymmetrical structures, such as quasimetrics. Extending VIP with recent method (Wang & Isola, 2022) that can learn quasimetrics with finite data may be a fruitful future direction.

We have also used VIP only as a frozen visual reward and representation module to test its broad generalization capability. Better absolute task performance may be achieved by fine-tuning VIP on task-specific data. Exploring how to best fine-tune VIP is a promising direction for pushing VIP’s limit.

Finally, we have focused on robot manipulation tasks in this work, but VIP’s training objective can also be used for pre-training reward and representation for other goal-directed tasks, such as visual navigation (Savva et al., 2019). Exploring how VIP can be used to solve these other embodied AI tasks is also a promising avenue of future work.