Q-Transformer: Scalable Offline Reinforcement Learning via Autoregressive Q-Functions

Yevgen Chebotar, Quan Vuong, Alex Irpan, Karol Hausman, Fei Xia, Yao Lu, Aviral Kumar, Tianhe Yu, Alexander Herzog, Karl Pertsch, Keerthana Gopalakrishnan, Julian Ibarz, Ofir Nachum, Sumedh Sontakke, Grecia Salazar, Huong T Tran, Jodilyn Peralta, Clayton Tan, Deeksha Manjunath, Jaspiar Singht, Brianna Zitkovich, Tomas Jackson, Kanishka Rao, Chelsea Finn, Sergey Levine

Introduction

Robotic learning methods that incorporate large and diverse datasets in combination with high-capacity expressive models, such as Transformers , have the potential to acquire generalizable and broadly applicable policies that perform well on a wide variety of tasks . For example, these policies can follow natural language instructions , perform multi-stage behaviors , and generalize broadly across environments, objects, and even robot morphologies . However, many of the recently proposed high-capacity models in the robotic learning literature are trained with supervised learning methods. As such, the performance of the resulting policy is limited by the degree to which human demonstrators can provide high-quality demonstration data. This is limiting for two reasons. First, we would like robotic systems that are more proficient than human teleoperators, exploiting the full potential of the hardware to perform tasks quickly, fluently, and reliably. Second, we would like robotic systems that get better with autonomously gathered experience, rather than relying entirely on high-quality demonstrations.

Reinforcement learning in principle provides both of these capabilities. A number of promising recent advances demonstrate the successes of large-scale robotic RL in varied settings, such as robotic grasping and stacking , learning heterogeneous tasks with human-specified rewards , learning multi-task policies , learning goal-conditioned policies , and robotic navigation . However, training high-capacity models such as Transformers using RL algorithms has proven more difficult to instantiate effectively at large scale. In this paper, we aim to combine large-scale robotic learning from diverse real-world datasets with modern high-capacity Transformer-based policy architectures.

While in principle simply replacing existing architectures (e.g., ResNets or smaller convolutional neural networks ) with a Transformer is conceptually straightforward, devising a methodology that effectively makes use of such architectures is considerably more challenging. High-capacity models only make sense when we train on large and diverse datasets – small, narrow datasets simply do not require this much capacity and do not benefit from it. While prior works used simulation to create such datasets , the most representative data comes from the real world . Therefore, we focus on reinforcement learning methods that can use Transformers and incorporate large, previously collected datasets via offline RL. Offline RL methods train on prior data, aiming to derive the most effective possible policy from a given dataset. Of course, this dataset can be augmented with additionally autonomously gathered data, but the training is separated from data collection, providing an appealing workflow for large-scale robotics applications .

Another issue in applying Transformer models to RL is to design RL systems that can effectively train such models. Effective offline RL methods generally employ Q-function estimation via temporal difference updates . Since Transformers model discrete token sequences, we convert the Q-function estimation problem into a discrete token sequence modeling problem, and devise a suitable loss function for each token in the sequence. Naïvely discretizing the action space leads to exponential blowup in action cardinality, so we employ a per-dimension discretization scheme, where each dimension of the action space is treated as a separate time step for RL. Different bins in the discretization corresponds to distinct actions. The per-dimension discretization scheme allows us to use simple discrete-action Q-learning methods with a conservative regularizer to handle distributional shift . We propose a specific regularizer that minimizes values of every action that was not taken in the dataset and show that our method can learn from both narrow demonstration-like data and broader data with exploration noise. Finally, we utilize a hybrid update that combines Monte Carlo and nn-step returns with temporal difference backups , and show that doing so improves the performance of our Transformer-based offline RL method on large-scale robotic learning problems.

In summary, our main contribution is the Q-Transformer, a Transformer-based architecture for robotic offline reinforcement learning that makes use of per-dimension tokenization of Q-values and can readily be applied to large and diverse robotic datasets, including real-world data. We summarize the components of Q-Transformer in Figure 1. Our experimental evaluation validates the Q-Transformer by learning large-scale text-conditioned multi-task policies, both in simulation for rigorous comparisons and in large-scale real-world experiments for realistic validation. Our real-world experiments utilize a dataset with 38,000 successful demonstrations and 20,000 failed autonomously collected episodes on more than 700 tasks, gathered with a fleet of 13 robots. Q-Transformer outperforms previously proposed architectures for large-scale robotic RL , as well as previously proposed Transformer-based models such as the Decision Transformer .

Related Work

Offline RL has been extensively studied in recent works . Conservative Q-learning (CQL) learns policies constrained to a conservative lower bound of the value function. Our goal is not to develop a new algorithmic principle for offline RL, but to devise an offline RL system that can integrate with high-capacity Transformers, and scale to real-world multi-task robotic learning. We thus develop a version of CQL particularly effective for training large Transformer-based Q-functions on mixed quality data. While some works have noted that imitation learning outperforms offline RL on demonstration data , other works showed offline RL techniques to be effective with demonstrations both in theory and in practice . Nonetheless, a setting that combines “narrow” demonstration data with “broad” sub-optimal (e.g., autonomously collected) data is known to be particularly difficult , though it is quite natural in many robotic learning settings where we might want to augment a core set of demonstrations with relatively inexpensive low-quality autonomously collected data. We believe that the effectiveness of our method in this setting is of particular interest to practitioners.

Transformer-based architectures have been explored in recent robotics research, both to learn generalizable task spaces and to learn multi-task or even multi-domain sequential policies directly . Although most of these works considered Transformers in a supervised learning setting, e.g., learning from demonstrations , there are works on employing Transformers for RL and conditional imitation learning . In our experiments, we compare to Decision Transformer (DT) in particular , which extends conditional imitation learning with reward conditioning to use sequence models, and structurally resembles imitation learning methods that have been used successfully for robotic control. Although DT incorporates elements of RL (namely, reward functions), it does not provide a mechanism to improve over the demonstrated behavior or recombine parts of the dataset to synthesize more optimal behaviors, and indeed is known to have theoretical limitations . On the other hand, such imitation-based recipes are popular perhaps due to the difficulty of integrating Transformer architectures with more powerful temporal difference methods (e.g., Q-learning). We show that several simple but important design decisions are needed to make this work, and our method significantly outperforms non-TD methods such as DT, as well as imitation learning, on our large-scale multi-task robotic control evaluation. Extending Decision Transformer, Yamagata et al. proposed to use a Q-function in combination with a Transformer-based policy, but the Q-function itself did not use a Transformer-based architecture. Our Q-function could in principle be combined with this method, but our focus is specifically on directly training Transformers to represent Q-values.

To develop a Transformer-based Q-learning method, we discretize each action space dimension, with each dimension acting as a distinct time step. Autoregressive generation of discrete actions has been explored by Metz et al. , who propose a hierarchical decomposition of an MDP and then utilize LSTM for autoregressive discretization. Our discretization scheme is similar but simpler, in that we do not use any hierarchical decomposition but simply treat each dimension as a time step. However, since our goal is to perform offline RL at scale with real-world image based tasks (vs. the smaller state-space tasks learned via online RL by Metz et al. ), we present a number of additional design decisions to impose a conservative regularizer, enabling training our Transformer-based offline Q-learning method at scale, providing a complete robotic learning system.

Background

In RL, we learn policies π\pi that maximizes the expected total reward in a Markov decision process (MDP) with states ss, actions aa, discount factor γ∈(0,1]\gamma\in(0,1], transition function T(s′∣s,a)T(s^{\prime}|s,a) and a reward function R(s,a)R(s,a). Actions aa have dimensionality dAd_{\mathcal{A}}. Value-based RL approaches learn a Q-function Q(s,a)Q(s,a) representing the total discounted return ∑tγtR(st,at)\sum_{t}\gamma^{t}R(s_{t},a_{t}), with policy π(a∣s)=arg max⁡aQ(s,a)\pi(a|s)=\operatorname*{arg\,max}_{a}Q(s,a). The Q-function can be learned by iteratively applying the Bellman operator :

approximated via function approximation and sampling. The offline RL setting assumes access to an offline dataset of transitions or episodes, produced by some unknown behavior policy πβ(a∣s)\pi_{\beta}(a|s), but does not assume the ability to perform additional online interaction during training. This is appealing for real-world robotic learning, where on-policy data collection is time-consuming. Learning from offline datasets requires addressing distributional shift, since in general the action that maximizes Q(st+1,at+1)Q(s_{t+1},a_{t+1}) might lie outside of the data distribution. One approach to mitigate this is to add a conservative penalty that pushes down the Q-values Q(s,a)Q(s,a) for any action aa outside of the dataset, thus ensuring that the maximum value action is in-distribution.

In this work, we consider tasks with sparse rewards, where a binary reward R∈{0,1}R\in\{0,1\} (indicating success or failure) is assigned at the last time step of episodes. Although our method is not specific to this setting, such reward structure is common in robotic manipulation tasks that either succeed or fail on each episode, and can be particularly challenging for RL due to the lack of reward shaping.

Q-Transformer

In this section, we introduce Q-Transformer, an architecture for offline Q-learning with Transformer models, which is based on three main ingredients. First, we describe how we apply discretization and autoregression to enable TD-learning with Transformer architectures. Next, we introduce a particular conservative Q-function regularizer that enables learning from offline datasets. Lastly, we show how Monte Carlo and nn-step returns can be used to improve learning efficiency.

Using Transformers with Q-learning presents two challenges: (1) we must tokenize the inputs to effectively apply attention mechanisms, which requires discretizing the action space; (2) we must perform maximization of Q-values over discretized actions while avoiding the curse of dimensionality. Addressing these issues within the standard Q-learning framework requires new modeling decisions. The intuition behind our autoregressive Q-learning update is to treat each action dimension as essentially a separate time step. That way, we can discretize individual dimensions (1D quantities), rather than the entire action space, avoiding the curse of dimensionality. This can be viewed as a simplified version of the scheme proposed in , though we apply this to high-capacity Transformer models, extend it to the offline RL setting, and scale it up to real-world robotic learning.

Let τ=(s1,a1,…,sT,aT)\tau=(s_{1},a_{1},\dots,s_{T},a_{T}) be a trajectory of robotic experience of length TT from an offline dataset D\mathcal{D}. For a given time-step tt, and the corresponding action ata_{t} in the trajectory, we define a per-dimension view of the action ata_{t}. Let at1:ia^{1:i}_{t} denote the vector of action dimensions from the first dimension at1a^{1}_{t} until the ii-th dimension atia^{i}_{t}, where ii can range from 11 to the total number of action dimensions, that we denote as dAd_{\mathcal{A}}. Then, for a time window ww of state history, we define the Q-value of the action atia^{i}_{t} in the i−thi-th dimension using an autoregressive Q-function conditioned on states from this time window st−w:ts_{t-w:t} and previous action dimensions for the current time step at1:i−1a_{t}^{1:i-1}. To train the Q-function, we define a per-dimension Bellman update. For all dimensions i∈{1,…,dA}i\in\{1,\dots,d_{\mathcal{A}}\}:

The reward is only applied on the last dimension (second line in the equation), as we do not receive any reward before executing the whole action. In addition, we only discount Q-values between the time steps and keep discounting at 1.01.0 for all but the last dimension within each time step, to ensure the same discounting as in the original MDP. Figure 2 illustrates this process, where each yellow box represents the Q-target computation with additional conservatism and Monte Carlo returns described in the next subsections. It should be noted that by treating each action dimension as a time step for the Bellman update, we do not change the general optimization properties of Q-learning algorithms and the principle of the Bellman optimality still holds for a given MDP as we maximize over an action dimension given the optimality of all action dimensions in the future. We show that this approach provides a theoretically consistent way to optimize the original MDP in Appendix A, with a proof of convergence in the tabular setting in Appendix B.

2 Conservative Q-Learning with Transformers

Having defined a Bellman backup for running Q-learning with Transformers, we now develop a technique that enables learning from offline data, including human demonstrations and autonomously collected data. This typically requires addressing over-estimation due to the distributional shift, when the Q-function for the target value is queried at an action that differs from the one on which it was trained. Conservative Q-learning (CQL) minimizes the Q-function on out-of-distribution actions, which can result in Q-values that are significantly smaller than the minimal possible cumulative reward that can be attained in any trajectory. When dealing with sparse rewards R∈{0,1}R\in\{0,1\}, results in show that the Q-function regularized with a standard conservative objective can take on negative values, even though instantaneous rewards are all non-negative. This section presents a modified version of conservative Q-learning that addresses this issue in our problem setting.

3 Improving Learning Efficiency with Monte Carlo and n𝑛n-step Returns

When the dataset contains some good trajectories (e.g., demonstrations) and some suboptimal trajectories (e.g., autonomously collected trials), utilizing Monte Carlo return-to-go estimates to accelerate Q-learning can lead to significant performance improvements, as the Monte Carlo estimates along the better trajectories lead to much faster value propagation. This has also been observed in prior work . Based on this observation, we propose a simple improvement to Q-Transformer that we found to be quite effective in practice. The Monte Carlo return is defined by the cumulative reward within the offline trajectory τ\tau: MCt:T=∑j=tTγj−tR(sj,aj)\text{MC}_{t:T}=\sum_{j=t}^{T}\gamma^{j-t}R(s_{j},a_{j}). This matches the Q-value of the behavior policy πβ\pi_{\beta}, and since the optimal Q∗(s,a)Q^{*}(s,a) is larger than the Q-value for any other policy, we have Q∗(st,at)≥MCt:TQ^{*}(s_{t},a_{t})\geq\text{MC}_{t:T}. Since the Monte Carlo return is a lower bound of the optimal Q-function, we can augment the Bellman update to take the maximum between the MC-return and the current Q-value: max⁡(MCt:T,Q(st,at))\max\left(\text{MC}_{t:T},Q(s_{t},a_{t})\right), without changing what the Bellman update will converge to.

Although this does not change convergence, including this maximization speeds up learning (see Section 5.3). We present a hypothesis why this occurs. In practice, Q-values for final timesteps (sT,aT)(s_{T},a_{T}) are learned first and then propagated backwards in future gradient steps. It can take multiple gradients for the Q-value to propagate all the way to (s1,a1)(s_{1},a_{1}). The max⁡(MC,Q)\max(\text{MC},Q) allows us to apply useful gradients to Q(s1,a1)Q(s_{1},a_{1}) at the start of training before the Q-values have propagated.

In our experiments, we also notice that additionally employing nn-step returns over action dimensions can significantly help with the learning speed. We pick nn such that the final Q-value of the last dimension of the next time step is used as the Q-target. This is because we get a new state and reward only after inferring and executing the whole action as opposed to parts of it, meaning that intermediate rewards remain 0 all the way until the last action dimension. While this introduces bias to the Bellman backups, as is always the case with off-policy learning with nn-step returns, we find in our ablation study in Section 5.3 that the detrimental effects of this bias are small, while the speedup in training is significant. This is consistent with previously reported results . More details about our Transformer sequence model architecture (depicted in Figure 3) conservative Q-learning implementation, and the robot system can be found in Appendix D.

Experiments

In our experiments, we aim to answer the following questions: (1) Can Q-Transformer learn from a combination of demonstrations and sub-optimal data? (2) How does Q-Transformer compare to other methods? (3) How important are the specific design choices in Q-Transformer? (4) Can Q-Transformer be applied to large-scale real world robotic manipulation problems?

Training dataset. The offline data used in our experiments was collected with a fleet of 13 robots, and consists of a subset of the demonstration data described by Brohan et al. , combined with lower quality autonomously collected data. The demonstrations were collected via human teleoperation for over 700700 distinct tasks, each with a separate language description. We use a maximum of 100 demonstrations per task, for a total of about 38,000 demonstrations. All of these demonstrations succeed on their respective tasks and receive a reward of 1.0. The rest of the dataset was collected by running the robots autonomously, executing policies learned via behavioral cloning.

To ensure a fair comparison between Q-Transformer and imitation learning methods, we discard all successful episodes in the autonomously collected data when we train our method, to ensure that by including the autonomous data the Q-Transformer does not get to observe more successful trials than the imitation learning baselines. This leaves us with about 20,000 additional autonomously collected failed episodes, each with a reward of 0.0, for a dataset size of about 58,000 episodes. The episodes are on average 3535 time steps in length. Examples of the tasks are shown in Figure 4.

Performance evaluation. To evaluate how well Q-Transformer can perform when learning from real-world offline datasets while effectively incorporating autonomously collected failed episodes, we evaluate Q-Transformer on 7272 unique manipulation tasks, and a variety of different skills, such as “drawer pick and place”, “open and close drawer”, “move object near target”, each consisting of 1818, 77 and 4848 unique tasks instructions respectively to specify different object combinations and drawers. As such, the average success rate in Table 4 is the average over 7272 tasks.

Since each task in the training set only has a maximum of 100100 demonstrations, we observe from Figure 4 that an imitation learning algorithm like RT-1 , which also uses a similar Transformer architecture, struggles to obtain a good performance when learning from only the limited pool of successful robot demonstrations. Existing offline RL methods, such as IQL and a Transformer-based method such as Decision Transformer , can learn from both successful demonstrations and failed episodes, and show better performance compared to RT-1, though by a relatively small margin. Q-Transformer has the highest success rate and outperforms both the behavior cloning baseline (RT-1) and offline RL baselines (Decision Transformer, IQL), exceeding the average performance of the best-performing prior method by about 70%. This demonstrates that Q-Transformer can effectively improve upon human demonstrations using autonomously collected sub-optimal data.

Appendix 8 also shows that Q-Transformer can be successfully applied in combination with a recently proposed language task planner to perform both affordance estimation and robot action execution. Q-Transformer outperforms prior methods for planning and executing long-horizon tasks.

2 Benchmarking in simulation

In this section, we evaluate Q-Transformer on a challenging simulated offline RL task that require incorporating sub-optimal data to solve the task. In particular, we use a visual simulated picking task depicted in Figure 5, where we have a small amount of position controlled human demonstrations (∼\sim8% of the data). The demonstrations are replayed with noise to generate more trajectories (∼\sim92% of the data). Figure 5 shows a comparison to several offline algorithms, such as QT-Opt with CQL , IQL , AW-Opt , and Decision Transformer , along with RT-1 using Behavioral Cloning on demonstrations only. As we see, algorithms that can effectively perform TD-learning to combine optimal and sub-optimal data (such as Q-Transformer and QT-Opt) perform better than others. BC with RT-1 is not able to take advantage of sub-optimal data. Decision Transformer is trained on both demonstrations and sub-optimal data, but is not able to leverage the noisy data for policy improvement and does not end up performing as well as our method. Although IQL and AW-Opt perform TD-learning, the actor remains too close to the data and can not fully leverage the sub-optimal data. Q-Transformer is able to both bootstrap the policy from demonstrations and also quickly improve through propagating information with TD-learning. We also analyze the statistical significance of the results by training with multiple random seeds in Appendix F.

3 Ablations

We perform a series of ablations of our method design choices in simulation, with results presented in Figure 6 (left). First, we demonstrate that our choice of conservatism for Q-Transformer performs better than the standard CQL regularizer, which corresponds to a softmax layer on top of the Q-function outputs with a cross-entropy loss between the dataset action and the output of this softmax . This regularizer plays a similar role to the one we propose, decreasing the Q-values for out-of-distribution actions and staying closer to the behavior policy.

As we see in Figure 6 (left), performance with softmax conservatism drops to around the fraction of demonstration episodes (∼\sim8%). This suggests a collapse to the behavior policy as the conservatism penalty becomes too good at constraining to the behavior policy distribution. Due to the nature of the softmax, pushing Q-values down for unobserved actions also pushes Q-values up for the observed actions, and we theorize this makes it difficult to keep Q-values low for sub-optimal in-distribution actions that fail to achieve high reward. Next, we show that using conservatism is important. When removing conservatism entirely, we observe that performance collapses. Actions that are rare in the dataset will have overestimated Q-values, since they are not trained by the offline Q-learning procedure. The resulting overestimated values will propagate and collapse the entire Q-function, as described in prior work . Finally, we ablate the Monte-Carlo returns and again observe performance collapse. This demonstrates that adding information about the sampled future returns significantly helps in bootstrapping the training of large architectures such as Transformers.

We also ablate the choice of nn-step returns from the Section 4.3 on real robots and observe that using nn-step returns leads to a significantly faster training speed as measured by the number of gradient steps and wall clock time compared to using 1-step returns, with a minimal loss in performance, as shown in Figure 6 (top right).

4 Massively scaling up Q-Transformer

The experiments in the previous section used a large dataset that included successful demonstrations and failed autonomous trials, comparable in size to some of the largest prior experiments that utilized demonstration data . We also carry out a preliminary experiment with a much larger dataset to investigate the performance of Q-Transformer as we scale up the dataset size.

This experiment includes all of the data collected with 13 robots and comprises of the demonstrations used by RT-1 and successful autonomous episodes, corresponding to about 115,000 successful trials, and an additional 185,000 failed autonomous episodes, for a total dataset size of about 300,000 trials. Model architecture and hyperparameters were kept exactly the same, as the computational cost of the experiment made further hyperparameter tuning prohibitive (in fact, we only train the models once). Note that with this number of successful demonstrations, even standard imitation learning with the RT-1 architecture already performs very well, attaining 82% success rate. However, as shown in Figure 6 (bottom right), Q-Transformer was able to improve even on this very high number. This experiment demonstrates that Q-Transformer can continue to scale to extremely large dataset sizes, and continues to outperform both imitation learning with RT-1 and Decision Transformer.

Limitations and Discussion

In this paper, we introduced the Q-Transformer, an architecture for offline reinforcement learning with high-capacity Transformer models that is suitable for large-scale multi-task robotic RL. Our framework does have several limitations. First, we focus on sparse binary reward tasks corresponding to success or failure for each trial. While this setup is reasonable for a broad range of episodic robotic manipulation problems, it is not universal, and we expect that Q-Transformer could be extended to more general settings as well in the future.

Second, the per-dimension action discretization scheme that we employ may become more cumbersome in higher dimensions (e.g., controlling a humanoid robot), as the sequence length and inference time for our model increases with action dimensionality. Although nn-step returns mitigate this to a degree, the length of the sequences still increases with action dimensionality. For such higher-dimensional action space, adaptive discretization methods might also be employed, for example by training a discrete autoencoder model and reducing representation dimensionality. Uniform action discretization can also pose problems for manipulation tasks that require a large range of motion granularities, e.g. both coarse and fine movements. In this case, adaptive discretization based on the distribution of actions could be used for representing both types of motions.

Finally, in this work we concentrated on the offline RL setting. However, extending Q-Transformer to online finetuning is an exciting direction for future work that would enable even more effective autonomous improvement of complex robotic policies.

References

Appendix A Proof of MDP optimization consistency

To show that transforming MDP into a per-action-dimension form still ensures optimization of the original MDP, we show that optimizing the Q-function for each action dimension is equivalent to optimizing the Q-function for the full action.

If we consider the full action a1:dAa_{1:d_{\mathcal{A}}} and that we switch to the state s′s^{\prime} at the next timestep, the Q-function for optimizing over the full action MDP would be:

where R(s,a1:dA∗)R(s,a_{1:d_{\mathcal{A}}}^{*}) is the reward we get after executing the full action.

The optimization over each action dimension using our Bellman update is:

which optimizes the original full action MDP as in Eq. 3.

Appendix B Proof of convergence

Convergence of Q-learning has been shown in the past . Below we demonstrate that per-action dimension Q-function converges as well, by providing a proof almost identical to the standard Q-learning convergence proof, but extended to account for the per-action dimension maximization.

Let dAd_{\mathcal{A}} be the dimensionality of the action space, aa indicates a possible sequence of actions, whose dimension is not necessarily equal to the dimension of the action space. That is:

To proof convergence, we can demonstrate that the Bellman operator applied to the per-action dimension Q-function is a contraction, i.e.:

a′a^{\prime} is the next action dimension following the sequence aa, s′s^{\prime} is the next state of the MDP, γ\gamma is the discounting factor, and 0≤c≤10\leq c\leq 1.

Proof: We can show that this is the case as follows:

Case 1: For action sequence whose dimension is less than the dimension of the action space.

where sup⁡s,a\sup_{s,a} is the supremum over all action sequences, with 0≤γ≤10\leq\gamma\leq 1 and ∣∣f∣∣∞=sup⁡x[f(x)]||f||_{\infty}=\sup_{x}[f(x)].

Case 2: For action sequence whose dimension is equal to the dimension of the action space

Appendix C Analysis of the conservatism term

With the goal of understanding the behavior of our training procedure, we theoretically analyze the solution obtained by Eq. 2 for the simpler cases when QQ is represented as a table, and when the objective in Eq. 2 can be minimized exactly. We derive the minimizer of the objective in Eq. 2 by differentiating JJ with respect to QQ:

Eq. 4 implies that training with the objective in Eq. 2 performs a weighted Bellman backup: unlike the standard Bellman backup, training with Eq. 2 multiplies large Q-value targets by a weight m(s,a)m(s,a). This weight m(s,a)m(s,a) takes values between 0 and 1, with larger values close to 1 for in-distribution actions where (s,a)∈D(s,a)\in\mathcal{D}, and very small values close to 0 for out-of-distribution actions aa at any state ss (i.e., actions where πβ(a∣s)\pi_{\beta}(a|s) is small). Thus, the Bellman backup induced via Eq. 4 should effectively prevent over-estimation of Q-values for unseen actions.

Appendix D Q-Transformer Architecture & System

In this section, we describe the architecture of Q-Transformer as well as the important implementation and system details that make it an effective Q-learning algorithm for real robots.

Our neural network architecture is shown in Figure 3. The architecture is derived from the RT-1 design , adapted to accommodate the Q-Transformer framework, and consists of a Transformer backbone that reads in images via a convolutional encoder followed by tokenization. Since we apply Q-Transformer to a multi-task robotic manipulation problem where each task is specified by a natural language instruction, we first embed the natural language instruction into an embedding vector via the Universal Sentence Encoder . The embedding vector and images from the robot camera are then converted into a sequence of input tokens via a FiLM EfficientNet . In the standard RT-1 architecture , the robot action space is discretized and the Transformer sequence model outputs the logits for the discrete action bins per dimension and per time step. In this work, we extend the network architecture to use Q-learning by applying a sigmoid activation to the output values for each action, and interpreting the resulting output after the sigmoid as Q-values. This representation is particularly suitable for tasks with sparse per-episode rewards R∈R\in, since the Q-values may be interpreted as probabilities of task success and should always lie in the range .Notethatunlikethestandardsoftmax,thisinterpretationofQ−valuesdoesnotprescribenormalizingacrossactions(i.e.,eachactionoutputcantakeonanyvaluein. Note that unlike the standard softmax, this interpretation of Q-values does not prescribe normalizing across actions (i.e., each action output can take on any value in).

Since our robotic system, described in Section D.3, has 8-dimensional actions, we end up with 8 dimensions per time step and discretize each one into N=256N=256 value bins. Our reward function is a sparse reward that assigns value 1.0 at the last step of an episode if the episode is successful and 0.0 otherwise. We use a discount rate γ=0.98\gamma=0.98. As is common in deep RL, we use a target network to estimate target Q-values QkQ^{k}, using an exponential moving average of QQ-network weights to update the target network. The averaging constant is set to 0.010.01.

D.2 Conservative Q-learning implementation

D.3 Robot system overview

The robot that we use in this work is a mobile manipulator with a 7-DOF arm with a 2 jaw parallel gripper, attached to a mobile base with a head-mounted RGB camera, illustrated in Figure 1. The RGB camera provides a 640×512640\times 512 RGB image, which is downsampled to 320×256320\times 256 before being consumed by the Q-Transformer. See Figure 4 for images from the robot camera view. The learned policy is set up to control the arm and the gripper of the robot. Our action space consists of 8 dimensions: 3D position, 3D orientation, gripper closure command, and an additional dimension indicating whether the episode should terminate, which the policy must trigger to receive a positive reward upon successful task completion. Position and orientation are relative to the current pose, while the gripper command is the absolute closedness fraction, ranging from fully open to fully closed. Orientation is represented via axis-angles, and all actions except whether to terminate are continuous actions discretized over their full action range in 256 bins. The termination action is binary, but we pad it to be the same size as the other action dimensions to avoid any issues with unequal weights. The policy operates at 3 Hz, with actions executed asynchronously .

Appendix E Pseudo-code

Algorithm 1 shows the loss computation for training each action dimension of the Q-Transformer. We first use Eq. 1 to compute the maximum Q-values over the next action dimensions. Then we compute the Q-target for the given dataset action by using the Bellman update with an additional maximization over the Monte-Carlo return and predicted maximum Q-value at the next time step. The TD-error is then computed using the Mean-Squared Error. Finally, we set a target of 0 for all discretized action bins except the dataset action and add the averaged Mean-Squared Error over these dimensions to the TD-Error, which results in the total loss L\mathcal{L}.

Appendix F Running training for multiple random seeds

In addition to performing a large amount of evaluations, we also analyze the statistical significance of our learning results by running our training of Q-Transformer and RT-1 on multiple seeds in simulation. In particular, we run the training for 5 random seeds in Figure 7. As we can see, Q-Transformer retains its improved performance across the distribution of the random seeds.

Appendix G Q-Transformer value function with a language planner experiments

Recently, the SayCan algorithm was proposed as a way to combine large language models (LLMs) with learned policies and value functions to solve long-horizon tasks. In this framework, the value function for each available skill is used to determine the “affordance” of the current state for that skill, and a large language model then selects from among the available affordances to take a step towards performing some temporally extended task. For example, if the robot is commanded to bring all the items on a table, the LLM might propose a variety of semantically meaningful items, and select from among them based on the item grasping skill that currently has a high value (corresponding to items that the robot thinks it can grasp). SayCan uses QT-Opt in combination with sim-to-real transfer to train Q-functions for these affordances. In the following set of experiments, we demonstrate that the Q-Transformer outperforms QT-Opt for affordance estimation without using any sim-to-real transfer, entirely using the real world dataset that we employ in the preceding experiments.

We first benchmark Q-Transformer on the problem of correctly estimating task affordances from the RT-1 dataset . In addition to the standard training on demonstrations and autonomous data, we introduce a training with relabeling, which we found particularly useful for affordance estimation. During relabeling, we sample a random alternate task for a given episode. We relabel the task name of the episode to the newly sampled task, and set reward to 0.0. This ensures that the boundaries between tasks are more clearly learned during training. Table 1 shows comparison of performance of our model with and without relabeling as well as the sim-to-real QT-Opt model used in SayCan . Both of our models outperform the QT-Opt model on F1 score, with the relabeled model outperforming it by a large margin. This demonstrates that our Q-function can be effectively used for affordance estimation, even without training with sim-to-real transfer. Visualization of the Q-values produced by our Q-function can be found in Figure 8.

We then use Q-Transformer in a long horizon SayCan style evaluation, replacing both the sim-to-real QT-Opt model for affordance estimation, and the RT-1 policy for low-level robotic control. During this evaluation, a PaLM language model is used to propose task candidates given a user query. Q-values are then used to pick the task candidate with the highest affordance score, which is then executed on the robot using the execution policy. The Q-Transformer used for affordance estimation is trained with relabeling. The Q-Transformer used for low-level control is trained without relabeling, since we found relabeling episodes at the task level did not improve execution performance. SayCan with Q-Transformer is better at both planning the sequence of tasks and executing those plans, as illustrated in Table 2.

Appendix H Real robotic manipulation tasks used in our evaluation

We include the complete list of evaluation tasks in our real robot experiments below.

Drawer pick and place: pick 7up can from top drawer and place on counter, place 7up can into top drawer, pick brown chip bag from top drawer and place on counter, place brown chip bag into top drawer, pick orange can from top drawer and place on counter, place orange can into top drawer, pick coke can from middle drawer and place on counter, place coke can into middle drawer, pick orange from middle drawer and place on counter, place orange into middle drawer, pick green rice chip bag from middle drawer and place on counter, place green rice chip bag into middle drawer, pick blue plastic bottle from bottom drawer and place on counter, place blue plastic bottle into bottom drawer, pick water bottle from bottom drawer and place on counter, place water bottle into bottom drawer, pick rxbar blueberry from bottom drawer and place on counter, place rxbar blueberry into bottom drawer.

Open and close drawer: open top drawer, close top drawer, open middle drawer, close middle drawer, open bottom drawer, close bottom drawer.

Move object near target: move 7up can near apple, move 7up can near blue chip bag, move apple near blue chip bag, move apple near 7up can, move blue chip bag near 7up can, move blue chip bag near apple, move blue plastic bottle near pepsi can, move blue plastic bottle near orange, move pepsi can near orange, move pepsi can near blue plastic bottle, move orange near blue plastic bottle, move orange near pepsi can, move redbull can near rxbar blueberry, move redbull can near water bottle, move rxbar blueberry near water bottle, move rxbar blueberry near redbull can, move water bottle near redbull can, move water bottle near rxbar blueberry, move brown chip bag near coke can, move brown chip bag near green can, move coke can near green can, move coke can near brown chip bag, move green can near brown chip bag, move green can near coke can, move green jalapeno chip bag near green rice chip bag, move green jalapeno chip bag near orange can, move green rice chip bag near orange can, move green rice chip bag near green jalapeno chip bag, move orange can near green jalapeno chip bag, move orange can near green rice chip bag, move redbull can near sponge, move sponge near water bottle, move sponge near redbull can, move water bottle near sponge, move 7up can near blue blastic bottle, move 7up can near green can, move blue plastic bottle near green can, move blue plastic bottle near 7up can, move green can near 7up can, move green can near blue plastic bottle, move apple near brown chip bag, move apple near green jalapeno chip bag, move brown chip bag near green jalapeno chip bag, move brown chip bag near apple, move green jalapeno chip bag near apple, move green jalapeno chip bag near brown chip bag.