Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning

Juan Rocamonde, Victoriano Montesinos, Elvis Nava, Ethan Perez, David Lindner

Introduction

Training reinforcement learning (RL) agents to perform complex tasks in vision-based domains can be difficult, due to high costs associated with reward specification. Manually specifying reward functions for real world tasks is often infeasible, and learning a reward model from human feedback is typically expensive. To make RL more useful in practical applications, it is critical to find a more sample-efficient and natural way to specify reward functions.

One natural approach is to use pretrained vision-language models (VLMs), such as CLIP (Radford et al., 2021) and Flamingo (Alayrac et al., 2022), to provide reward signals based on natural language. However, prior attempts to use VLMs to provide rewards require extensive fine-tuning VLMs (e.g., Du et al., 2023) or complex ad-hoc procedures to extract rewards from VLMs (e.g., Mahmoudieh et al., 2022). In this work, we demonstrate that simple techniques for using VLMs as zero-shot language-grounded reward models work well, as long as the chosen underlying model is sufficiently capable. Concretely, we make four key contributions.

First, we propose VLM-RM, a general method for using pre-trained VLMs as a reward model for vision-based RL tasks (Section 3). We propose a concrete implementation that uses CLIP as a VLM and cos-similarity between the CLIP embedding of the current environment state and a simple language prompt as a reward function. We can optionally regularize the reward model by providing a “baseline prompt” that describes a neutral state of the environment and partially projecting the representations onto the direction between baseline and target prompts when computing the reward.

Second, we validate our method in the standard CartPole and MountainCar RL benchmarks (Section 4.2). We observe high correlation between VLM-RMs and the ground truth rewards of the environments and successfully train policies to solve the tasks using CLIP as a reward model. Furthermore, we find that the quality of CLIP as a reward model improves if we render the environment using more realistic textures.

Third, we train a MuJoCo humanoid to learn complex tasks, including raising its arms, sitting in a lotus position, doing the splits, and kneeling (Figure 1; Section 4.3) using a CLIP reward model derived from single sentence text prompts (e.g., “a humanoid robot kneeling”).

Fourth, we study how VLM-RMs’ performance scales with the size of the VLM, and find that VLM scale is strongly correlated to VLM-RM quality (Section 4.4). In particular, we can only learn the humanoid tasks in Figure 1 with the largest publicly available CLIP model.

Our results indicate that VLMs are powerful zero-shot reward models. While current models, such as CLIP, have important limitations that persist when used as VLM-RMs, we expect such limitations to mostly be overcome as larger and more capable VLMs become available. Overall, VLM-RMs are likely to enable us to train models to perform increasingly sophisticated tasks from human-written task descriptions.

Background

At each point in time, the environment is in a state s∈Ss\in\mathcal{S}. In each timestep, the agent takes an action a∈Aa\in\mathcal{A}, causing the environment to transition to state s′s^{\prime} with probability θ(s′∣s,a)\theta(s^{\prime}|s,a). The agent then receives an observation oo, with probability ϕ(o∣s′)\phi(o|s^{\prime}) and a reward r=R(s,a,s′)r=R(s,a,s^{\prime}). A sequence of states and actions is called a trajectory τ=(s0,a0,s1,a1,… )\tau=(s_{0},a_{0},s_{1},a_{1},\dots), where si∈Ss_{i}\in\mathcal{S}, and ai∈Aa_{i}\in\mathcal{A}. The returns of such a trajectory τ\tau are the discounted sum of rewards g(τ;R)=∑t=0γtR(st,at,st+1)g(\tau;R)=\sum_{t=0}\gamma^{t}R(s_{t},a_{t},s_{t+1}).

We broadly define vision-language models (VLMs; Zhang et al., 2023) as models capable of processing sequences of both language inputs l∈L≤nl\in\mathcal{L}^{\leq n} and vision inputs i∈I≤mi\in\mathcal{I}^{\leq m}. Here, L\mathcal{L} is a finite alphabet and L≤n\mathcal{L}^{\leq n} contains strings of length less than or equal to nn, whereas I\mathcal{I} is the space of 2D RGB images and I≤m\mathcal{I}^{\leq m} contains sequences of images with length less than or equal to mm.

Vision-Language Models as Reward Models (VLM-RMs)

This section presents how we can use VLMs as a learning-free (zero-shot) way to specify rewards from natural language descriptions of tasks. Importantly, VLM-RMs avoid manually engineering a reward function or collecting expensive data for learning a reward model.

Let us consider a POMDP without a reward function (S,A,θ,O,ϕ,γ,d0)(\mathcal{S},\mathcal{A},\theta,\mathcal{O},\phi,\gamma,d_{0}). We focus on vision-based RL where the observations o∈Oo\in\mathcal{O} are images. For simplicity, we assume a deterministic observation distribution ϕ(o∣s)\phi(o|s) defined by a mapping ψ(s):S→O\psi(s):\mathcal{S}\rightarrow\mathcal{O} from states to image observation. We want the agent to perform a task T\mathcal{T} based on a natural language description l∈L≤nl\in\mathcal{L}^{\leq n}. For example, when controlling a humanoid robot (Section 4.3) T\mathcal{T} might be the robot kneeling on the ground and \l\l might be the string “a humanoid robot kneeling”.

To train the agent using RL, we need to first design a reward function. We propose to use a VLM to provide the reward R(s)R(s) as:

where c∈L≤nc\in\mathcal{L}^{\leq n} is an optional context, e.g., for defining the reward interactively with a VLM. This formulation is general enough to encompass the use of several different kinds of VLMs, including image and video encoders, as reward models.

In our experiments, we chose a CLIP encoder as the VLM. A very basic way to use CLIP to define a reward function is to use cosine similarity between a state’s image representation and the natural language task description:

In this case, we do not require a context cc. We will sometimes call the CLIP image encoder a state encoder, as it encodes an image that is a direct function of the POMDP state, and the CLIP language encoder a task encoder, as it encodes the language description of the task.

2 Goal-Baseline Regularization to Improve CLIP Reward Models

While in the previous section, we introduced a very basic way of using CLIP to define a task-based reward function, this section proposes Goal-Baseline Regularization as a way to improve the quality of the reward by projecting out irrelevant information about the observation.

So far, we assumed we only have a task description l∈L≤nl\in\mathcal{L}^{\leq n}. To apply goal-baseline regularization, we require a second “baseline” description b∈L≤nb\in\mathcal{L}^{\leq n}. The baseline bb is a natural language description of the environment setting in its default state, irrespective of the goal. For example, our baseline description for the humanoid is simply “a humanoid robot,” whereas the task description is, e.g., “a humanoid robot kneeling.” We obtain the goal-baseline regularized CLIP reward model (RCLIP-RegR_{\text{CLIP-Reg}}) by projecting our state embedding onto the line spanned by the baseline and task embeddings.

Given a goal task description ll and baseline description bb, let g=CLIPL(l)∥CLIPL(l)∥\mathbf{g}=\frac{\text{CLIP}_{L}(l)}{\|\text{CLIP}_{L}(l)\|}, b=CLIPL(b)∥CLIPL(b)∥\mathbf{b}=\frac{\text{CLIP}_{L}(b)}{\|\text{CLIP}_{L}(b)\|}, s=CLIPI(ψ(s))∥CLIPI(ψ(s))∥\mathbf{s}=\frac{\text{CLIP}_{I}(\psi(s))}{\|\text{CLIP}_{I}(\psi(s))\|} be the normalized encodings, and LL be the line spanned by b\mathbf{b} and g\mathbf{g}. The goal-baseline regularized reward function is given by

where α\alpha is a parameter to control the regularization strength.

In particular, for α=0\alpha=0, we recover our initial CLIP reward function RCLIPR_{\text{CLIP}}. On the other hand, for α=1\alpha=1, the projection removes all components of s\mathbf{s} orthogonal to g−b\mathbf{g}-\mathbf{b}.

Intuitively, the direction from b\mathbf{b} to g\mathbf{g} captures the change from the environment’s baseline to the target state. By projecting the reward onto this direction, we directionally remove irrelevant parts of the CLIP representation. However, we can not be sure that the direction really captures all relevant information. Therefore, instead of using α=1\alpha=1, we treat it as a hyperparameter. However, we find the method to be relatively robust to changes in α\alpha with most intermediate values being better than or 11.

3 RL with CLIP Reward Model

We can now use VLM-RMs as a drop-in replacement for the reward signal in RL. In our implementation, we use the Deep Q-Network (DQN; Mnih et al., 2015) or Soft Actor-Critic (SAC; Haarnoja et al., 2018) RL algorithms. Whenever we interact with the environment, we store the observations in a replay buffer. In regular intervals, we pass a batch of observations from the replay buffer through a CLIP encoder to obtain the corresponding state embeddings. We can then compute the reward function as cosine similarity between the state embeddings and the task embedding which we only need to compute once. Once we have computed the reward for a batch of interactions, we can use them to perform the standard RL algorithm updates. Appendix C contains more implementation details and pseudocode for our full algorithm in the case of SAC.

Experiments

We conduct a variety of experiments to evaluate CLIP as a reward model with and without goal-baseline regularization. We start with simple control tasks that are popular RL benchmarks: CartPole and MountainCar (Section 4.2). These environments have a ground truth reward function and a simple, well-structured state space. We find that our reward models are highly correlated with the ground truth reward function, with this correlation being greatest when applying goal-baseline regularization. Furthermore, we find that the reward model’s outputs can be significantly improved by making a simple modification to make the environment’s observation function more realistic, e.g., by rendering the mountain car over a mountain texture.

We then move on to our main experiment: controlling a simulated humanoid robot (Section 4.3). We use CLIP reward models to specify tasks from short language prompts; several of these tasks are challenging to specify manually. We find that these zero-shot CLIP reward models are sufficient for RL algorithms to learn most tasks we attempted with little to no prompt engineering or hyperparameter tuning.

Finally, we study the scaling properties of the reward models by using CLIP models of different sizes as reward models in the humanoid environment (Section 4.4). We find that larger CLIP models are significantly better reward models. In particular, we can only successfully learn the tasks presented in Figure 1 when using the largest publicly available CLIP model.

We extend the implementation of the DQN and SAC algorithm from the stable-baselines3 library (Raffin et al., 2021) to compute rewards from CLIP reward models instead of from the environment. As shown in Algorithm 1 for SAC, we alternate between environment steps, computing the CLIP reward, and RL algorithm updates. We run the RL algorithm updates on a single NVIDIA RTX A6000 GPU. The environment simulation runs on CPU, but we perform rendering and CLIP inference distributed over 4 NVIDIA RTX A6000 GPUs.

We provide the code to reproduce our experiments in the supplementary material. We discuss hyperparameter choices in Appendix C, but we mostly use standard parameters from stable-baselines3. Appendix C also contains a table with a full list of prompts for our experiments, including both goal and baseline prompts when using goal-baseline regularization.

1 How can we Evaluate VLM-RMs?

Evaluating reward models can be difficult, particularly for tasks for which we do not have a ground truth reward function. In our experiments, we use 3 types of evaluation: (i) evaluating policies using ground truth reward; (ii) comparing reward functions using EPIC distance; (iii) human evaluation.

If we have a ground truth reward function for a task such as for the CarPole and MountainCar, we can use it to evaluate policies. For example, we can train a policy using a VLM-RM and evaluate it using the ground truth reward. This is the most popular way to evaluate reward models in the literature and we use it for environments where we have a ground-truth reward available.

The “Equivalent Policy-Invariant Comparison” (EPIC; Gleave et al., 2021) distance compares two reward functions without requiring the expensive policy training step. EPIC distance is provably invariant on the equivalence class of reward functions that induce the same optimal policy. We consider only goal-based tasks, for which the EPIC is distance particularly easy to compute. In particular, a low EPIC distance between the CLIP reward model and the ground truth reward implies that the CLIP reward model successfully separates goal states from non-goal states. Appendix A discusses in more detail how we compute the EPIC distance in our case, and how we can intuitively interpret it for goal-based tasks.

For tasks without a ground truth reward function, such as all humanoid tasks in Figure 1, we need to perform human evaluations to decide whether our agent is successful. We define “success rate” as the percentage of trajectories in which the agent successfully performs the task in at least 50%50\% of the timesteps. For each trajectory, we have a single raterOne of the authors. label how many timesteps were spent successfully performing the goal task, and use this to compute the success rate. However, human evaluations can also be expensive, particularly if we want to evaluate many different policies, e.g., to perform ablations. For such cases, we additionally collect a dataset of human-labelled states for each task, including goal states and non-goal states. We can then compute the EPIC distance with these binary human labels. Empirically, we find this to be a useful proxy for the reward model quality which correlates well with the performance of a policy trained using the reward model.

For more details on our human evaluation protocol, we refer to Appendix B. Our human evaluation protocol is very basic and might be biased. Therefore, we additionally provide videos of our trained agents at https://sites.google.com/view/vlm-rm.

2 Can VLM-RMs Solve Classic Control Benchmarks?

As an initial validation of our methods, we consider two classic control environments: CartPole and MountainCar, implemented in OpenAI Gym (Brockman et al., 2016). In addition to the default MountainCar environment, we also consider a version with a modified rendering method that adds textures to the mountain and the car so that it resembles the setting of “a car at the peak of a mountain” more closely (see Figure 2). This environment allows us to test whether VLM-RMs work better in visually “more realistic” environments.

To understand the rewards our CLIP reward models provide, we first analyse plots of their reward landscape. In order to obtain a simple and interpretable visualization figure, we plot CLIP rewards against a one-dimensional state space parameter, that is directly related to the completion of the task. For the CartPole (Figure 2(a)) we plot CLIP rewards against the angle of the pole, where the ideal position is at angle . For the (untextured and textured) MountainCar environments Figures 2(b) and 2(c), we plot CLIP rewards against the position of the car along the horizontal axis, with the goal location being around x=0.5x=0.5.

Figure 2(a) shows that CLIP rewards are well-shaped around the goal state for the CartPole environment, whereas Figure 2(b) shows that CLIP rewards for the default MountainCar environment are poorly shaped, and might be difficult to learn from, despite still having roughly the right maximum.

We conjecture that zero-shot VLM-based rewards work better in environments that are more “photorealistic” because they are closer to the training distribution of the underlying VLM. Figure 2(c) shows that if, as described earlier, we apply custom textures to the MountainCar environment, the CLIP rewards become well-shaped when used in concert with the goal-baseline regularization technique. For larger regularization strength α\alpha, the reward shape resembles the slope of the hill from the environment itself – an encouraging result.

We then train agents using the CLIP rewards and goal-baseline regularization in all three environments, and achieve 100% task success rate in both environments (CartPole and textured MountainCar) for most α\alpha regularization strengths. Without the custom textures, we are not able to successfully train an agent on the mountain car task, which supports our hypothesis that the environment visualization is too abstract.

The results show that both and regularized CLIP rewards are effective in the toy RL task domain, with the important caveat that CLIP rewards are only meaningful and well-shaped for environments that are photorealistic enough for the CLIP visual encoder to interpret correctly.

3 Can VLM-RMs Learn Complex, Novel Tasks in a Humanoid Robot?

Our primary goal in using VLM-RMs is to learn tasks for which it is difficult to specify a reward function manually. To study such tasks, we consider the Humanoid-v4 environment implemented in the MuJoCo simulator (Todorov et al., 2012).

The standard task in this environment is for the humanoid robot to stand up. For this task, the environment provides a reward function based on the vertical position of the robot’s center of mass. We consider a range of additional tasks for which no ground truth reward function is available, including kneeling, sitting in a lotus position, and doing the splits. For a full list of tasks we tested, see Table 1. Appendix C presents more detailed task descriptions and the full prompts we used.

We make two modifications to the default Humanoid-v4 environment to make it better suited for our experiments. (1) We change the colors of the humanoid texture and the environment background to be more realistic (based on our results in Section 4.2 that suggest this should improve the CLIP encoder). (2) We move the camera to a fixed position pointing at the agent slightly angled down because the original camera position that moves with the agent can make some of our tasks impossible to evaluate. We ablate these changes in Figure 3, finding the texture change is critical and repositioning the camera provides a modest improvement.

Table 1 shows the human-evaluated success rate for all tasks we tested. We solve 5 out of 8 tasks we tried with minimal prompt engineering and tuning. For the remaining 3 tasks, we did not get major performance improvements with additional prompt engineering and hyperparameter tuning, and we hypothesize these failures are related to capability limitations in the CLIP model we use. We invite the reader to evaluate the performance of the trained agents themselves by viewing videos at https://sites.google.com/view/vlm-rm.

The three tasks that the agent does not obtain perfect performance for are “hands on hips”, “standing on one leg”, and “arms crossed”. We hypothesize that “standing on one leg” is very hard to learn or might even be impossible in the MuJoCo physics simulation because the humanoid’s feet are round. The goal state for “hands on hips” and “arms crossed” is visually similar to a humanoid standing and we conjecture the current generation of CLIP models are unable to discriminate between such subtle differences in body pose.

While the experiments in Table 1 use no goal-baseline regularization (i.e., α=0\alpha=0), we separately evaluate goal-baseline regularization for the kneeling task. Figure 4(a) shows that α≠0\alpha\neq 0 improves the reward model’s EPIC distance to human labels, suggesting that it would also improve performance on the final task, we might need a more fine-grained evaluation criterion to see that.

4 How do VLM-RMs Scale with VLM Model Size?

Finally, we investigate the effect of the scale of the pre-trained VLM on its quality as a reward model. We focus on the “kneeling” task and consider 4 different large CLIP models: the original CLIP RN50 (Radford et al., 2021), and the ViT-L-14, ViT-H-14, and ViT-bigG-14 from OpenCLIP (Cherti et al., 2023) trained on the LAION-5B dataset (Schuhmann et al., 2022).

In Figure 4(a) we evaluate the EPIC distance to human labels of CLIP reward models for the four model scales and different values of α\alpha, and we evaluate the success rate of agents trained using the four models. The results clearly show that VLM model scale is a key factor in obtaining good reward models. We detect a clear positive trend between model scale, and the EPIC distance of the reward model from human labels. On the models we evaluate, we find the EPIC distance to human labels is close to log-linear in the size of the CLIP model (Figure 4(b)).

This improvement in EPIC distance translates into an improvement in success rate. In particular, we observe a sharp phase transition between the ViT-H-14 and VIT-bigG-14 CLIP models: we can only learn the kneeling task successfully when using the VIT-bigG-14 model and obtain 0%0\% success rate for all smaller models (Figure 4(c)). Notably, the reward model improves smoothly and predictably with model scale as measured by EPIC distance. However, predicting the exact point where the RL agent can successfully learn the task is difficult. This is a common pattern in evaluating large foundation models, as observed by Ganguli et al. (2022).

Related Work

Foundation models (Bommasani et al., 2021) trained on large scale data can learn remarkably general and transferable representations of images, language, and other kinds of data, which makes them useful for a large variety of downstream tasks. For example, pre-trained vision-language encoders, such as CLIP (Radford et al., 2021), have been used far beyond their original scope, e.g., for image generation (Ramesh et al., 2022; Patashnik et al., 2021; Nichol et al., 2021), robot control (Shridhar et al., 2022; Khandelwal et al., 2022), or story evaluation (Matiana et al., 2021).

Reinforcement learning from human feedback (RLHF; Christiano et al., 2017) is a critical step in making foundation models more useful (Ouyang et al., 2022). However, collecting human feedback is expensive. Therefore, using pre-trained foundation models themselves to obtain reward signals for RL finetuning has recently emerged as a key paradigm in work on large language models (Bai et al., 2022). Some approaches only require a small amount of natural language feedback instead of a whole dataset of human preferences (Scheurer et al., 2022; 2023; Chen et al., 2023). However, similar techniques have yet to be adopted by the broader RL community.

While some work uses language models to compute a reward function from a structured environment representation (Xie et al., 2023), many RL tasks are visual and require using VLMs instead. Cui et al. (2022) use CLIP to provide rewards for robotic manipulation tasks given a goal image. However, they only show limited success when using natural language descriptions to define goals, which is the focus of our work. Mahmoudieh et al. (2022) are the first to successfully use CLIP encoders as a reward model conditioned on language task descriptions in robotic manipulation tasks. However, to achieve this, the authors need to explicitly fine-tune the CLIP image encoder on a carefully crafted dataset for a robotics task. Instead, we focus on leveraging CLIP’s zero-shot ability to specify reward functions, which is significantly more sample-efficient and practical. Du et al. (2023) finetune a Flamingo VLM (Alayrac et al., 2022) to act as a “success detector” for vision-based RL tasks tasks. However, they do not train RL policies using these success detectors, leaving open the question of how robust they are under optimization pressure.

In contrast to these works, we do not require any finetuning to use CLIP as a reward model, and we successfully train RL policies to achieve a range of complex tasks that do not have an easily-specified ground truth reward function.

Conclusion

We introduced a method to use vision-language models (VLMs) as reward models for reinforcement learning (RL), and implemented it using CLIP as a reward model and standard RL algorithms. We used VLM-RMs to solve classic RL benchmarks and to learn to perform complicated tasks using a simulated humanoid robot. We observed a strong scaling trend with model size, which suggests that future VLMs are likely to be useful as reward models in an even broader range of tasks.

Fundamentally, our approach relies on the reward model generalizing from a text description to a reward function that captures what a human intends the agent to do. Although the concrete failure cases we observed are likely specific to the CLIP models we used and may be solved by more capable models, some problems will persist. The resulting reward model will be misspecified if the text description does not contain enough information about what the human intends or the VLM generalizes poorly. While we expect future VLMs to generalize better, the risk of the reward model being misspecified grows for more complex tasks, that are difficult to specify in a single language prompt, and in practical applications with larger potential risks. Therefore, when using VLM-RMs in practice it will be crucial to use independent monitoring to ensure agents trained from automated feedback act as intended. For complex tasks, it will be prudent to use a multi-step reward specification, e.g., by using a VLM capable of having a dialogue with the user about specifying the task.

We were able to learn complex tasks using a simple approach to construct a reward model from CLIP. There are many possible extensions of our implementation that may be able to improve performance but were not necessary in our tasks. Finetuning VLMs for specific environments is a natural next step to make them more useful as reward models. To move beyond goal-based supervision, future VLM-RMs could use VLMs that can encode videos instead of images. To move towards specifying more complex tasks, future VLM-RMs could use dialogue-enabled VLMs.

For practical applications, it will be particularly important to ensure robustness and safety of the reward model. Our work can serve as a basis for studying the safety implications of VLM-RMs. For instance, future work could investigate the robustness of VLM-RMs against optimization pressure by RL agents and aim to identify instances of specification gaming.

More broadly, we believe VLM-RMs open up exciting avenues for future research to build useful agents on top of pre-trained models, such as building language model agents and real world robotic controllers for tasks where we do not have a reward function available.

Author Contributions

Juan Rocamonde designed and implemented the experimental infrastructure, ran most experiments, analyzed results, and wrote large parts of the paper.

Victoriano Montesinos implemented parallelized rendering and training to enable using larger CLIP models, implemented and ran many experiments, and performed the human evaluations.

Elvis Nava advised on experiment design, implemented and ran some of the experiments, and wrote large parts of the paper.

Ethan Perez proposed the original project and advised on research direction and experiment design.

David Lindner implemented and ran early experiments with the humanoid robot, wrote large parts of the paper, and led the project.

Acknowledgments

We thank Adam Gleave for valuable discussions throughout the project and detailed feedback on an early version of the paper, Jérémy Scheurer for helpful feedback early on, Adrià Garriga-Alonso for help with running experiments, and Xander Balwit for help with editing the paper.

We are grateful for funding received by Open Philanthropy, Manifund, the ETH AI Center, Swiss National Science Foundation (B.F.G. CRSII5-173721 and 315230 189251), ETH project funding (B.F.G. ETH-20 19-01), and the Human Frontiers Science Program (RGY0072/2019).

References

Appendix A Computing and Interpreting EPIC Distance

Our experiments all have goal-based ground truth reward functions, i.e., they give high reward if a goal state is reached and low reward if not. This section discusses how this helps to estimate EPIC distance between reward functions more easily. As a side-effect, this gives us an intuitive understanding of EPIC distance in our context. First, let us define EPIC distance.

The Equivalent-Policy Invariant Comparison (EPIC) distance between reward functions R1R_{1} and R2R_{2} is:

where ρ(⋅,⋅)\rho(\cdot,\cdot) is the Pearson correlation w.r.t a given distribution over transitions, and C(R)\mathcal{C}(R) is the canonically shaped reward, defined as:

where STC=S∖STS_{\mathcal{T}}^{C}=\mathcal{S}\setminus S_{\mathcal{T}} .

First, note that for reward functions where the reward of a transition (s,a,s′)(s,a,s^{\prime}) only depends on s′s^{\prime}, the canonically-shaped reward simplifies to:

Hence, because the Pearson correlation is location-invariant, we have

In practice, we use Lemma 1 to evaluate EPIC distance between a CLIP reward model and a ground truth reward function.

Note that the EPIC distance depends on a state distribution μ\mu (see Gleave et al. (2021) for further discussion). In our experiment, we use either a uniform distribution over states (for the toy RL environments) or the state distribution induced by a pre-trained expert policy (for the humanoid experiments). More details on how we collected the dataset for evaluating EPIC distances can be found in the Appendix B.

Appendix B Human Evaluation

Evaluation on tasks for which we do not have a reward function was done manually by one of the authors, depending on the amount of time the agent met the criteria listed in Table 2. See Figures 5 and 6 for the raw labels obtained about the agent performance.

We further evaluated the impact of goal-baseline regularization on the humanoid tasks that did not succeed in our experiments with α=0\alpha=0, cf. Figure 8. In these cases, goal baseline regularization does not improve performance. Together with the results in Figure 4(a), this could suggest that goal-baseline regularization is more useful for smaller CLIP models than for larger CLIP models. Alternatively, it is possible that the improvements to the reward model obtained by goal-baseline regularization are too small to lead to noticeable performance increases in the trained agents for the failing humanoid tasks. Unfortunately, a more thorough study of this was infeasible due to the cost associated with human evaluations.

Our second type of human evaluation is to compute the EPIC distance of a reward model to a pre-labelled set of states. To create a dataset for these evaluations, we select all checkpoints from the training run with the highest VLM-RM reward of the largest and most capable VLM we used. We then collect rollouts from each checkpoint and collect the images across all timesteps and rollouts into a single dataset. We then have a human labeller (again an author of this paper) label each image according to whether it represents the goal state or not, using the same criteria from Table 2. We use such a dataset for Figure 4. Figure 7 shows a more detailed breakdown of the EPIC distance for different model scales.

Appendix C Implementation Details & Hyperparameter Choices

In this section, we describe implementation details for both our toy RL environment experiments and the humanoid experiments, going into further detail on the experiment design, any modifications we make to the simulated environments, and the hyperparameters we choose for the RL algorithms we use.

Algorithm 1 shows pseudocode of how we integrate computing CLIP rewards with a batched RL algorithm, in this case SAC.

We use the standard CartPole and MountainCar environments implemented in Gym, but remove the termination conditions. Instead the agent receives a negative reward for dropping the pole in CartPole and a positive reward for reaching the goal position in MountainCar. We make this change because the termination leaks information about the task completion such that without removing the termination, for example, any positive reward function will lead to the agent solving the CartPole task. As a result of removing early termination conditions, we make the goal state in the MountainCar an absorbing state of the Markov process. This is to ensure that the estimated returns are not affected by anything a policy might do after reaching the goal state. Otherwise, this could, in particular, change the optimal policy or make evaluations much noisier.

We use DQN (Mnih et al., 2015) for CartPole, our only environment with a discrete action space, and SAC (Haarnoja et al., 2018), which is designed for continuous environments, for MountainCar. For both algorithms, we use a standard implementation provided by stable-baselines3 (Raffin et al., 2021).

We train for 33 million steps with a fixed episode length of 200200 steps, where we start the training after collecting 7500075000 steps. Every 200200 steps, we perform 200200 DQN updates with a learning rate of 2.3e−32.3e-3. We save a model checkpoint every 6400064000 steps. The Q-networks are represented by a 22 layer MLP of width 256256.

We train for 33 million steps using SAC parameters τ=0.01\tau=0.01, γ=0.9999\gamma=0.9999, learning rate 10−410^{-}4 and entropy coefficient 0.10.1. The policy is represented by a 22 layer MLP of width 6464. All other parameters have the default value provided by stable-baselines3.

We chose these hyperparameters in preliminary experiments with minimal tuning.

C.2 Humanoid Environment

For all humanoid experiments, we use SAC with the same set of hyperparameters tuned on preliminary experiments with the kneeling task. We train for 1010 million steps with an episode length of 100100 steps. Learning starts after 5000050000 initial steps and we do 100100 SAC updates every 100100 environment steps. We use SAC parameters τ=0.005\tau=0.005, γ=0.95\gamma=0.95, and learning rate 6⋅10−46\cdot 10^{-4}. We save a model checkpoint every 128000128000 steps. For our final evaluation, we always evaluate the checkpoint with the highest training reward. We parallelize rendering over 4 GPUs, and also use batch size B=3200B=3200 for evaluating the CLIP rewards.