Mastering Diverse Domains through World Models

Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, Timothy Lillicrap

Introduction

Reinforcement learning has enabled computers to solve individual tasks through interaction, such as surpassing humans in the games of Go and Dota 1, 2. However, applying algorithms to new application domains, for example from board games to video games or robotics tasks, requires expert knowledge and computational resources for tuning the algorithms 3. This brittleness also hinders scaling to large models that are expensive to tune. Different domains pose unique learning challenges that have prompted specialized algorithms, such as for continuous control 4, 5, sparse rewards 6, 7, image inputs 8, 9, and spatial environments 10, 11. Creating a general algorithm that learns to master new domains out of the box—without tuning—would overcome the barrier of expert knowledge and open up reinforcement learning to a wide range of practical applications.

We present DreamerV3, a general and scalable algorithm that masters a wide range of domains with fixed hyperparameters, outperforming specialized algorithms. DreamerV3 learns a world model 12, 13, 14 from experience for rich perception and imagination training. The algorithm consists of 3 neural networks: the world model predicts future outcomes of potential actions, the critic judges the value of each situation, and the actor learns to reach valuable situations. We enable learning across domains with fixed hyperparameters by transforming signal magnitudes and through robust normalization techniques. To provide practical guidelines for solving new challenges, we investigating the scaling behavior of DreamerV3. Notably, we demonstrate that increasing the model size of DreamerV3 monotonically improves both its final performance and data-efficiency.

The popular video game Minecraft has become a focal point of reinforcement learning research in recent years, with international competitions held for learning to collect diamonds in Minecraft 15. Solving this challenge without human data has been widely recognized as a milestone for artificial intelligence because of the sparse rewards, exploration difficulty, and long time horizons in this procedurally generated open-world environment. Due to these obstacles, previous approaches resorted to human expert data and manually-crafted curricula 16, 17. DreamerV3 is the first algorithm to collect diamonds in Minecraft from scratch, solving this challenge.

We summarize the four key contributions of this paper as follows:

We present DreamerV3, a general algorithm that learns to master diverse domains while using fixed hyperparameters, making reinforcement learning readily applicable.

We demonstrate the favorable scaling properties of DreamerV3, where increased model size leads to monotonic improvements in final performance and data-efficiency.

We perform an extensive evaluation, showing that DreamerV3 outperforms more specialized algorithms across domains, and release the training curves of all methods to facilitate comparison.

We find that DreamerV3 is the first algorithm to collect diamonds in Minecraft from scratch without human data or curricula, solving a long-standing challenge in artificial intelligence.

DreamerV3

The DreamerV3 algorithm consists of 3 neural networks—the world model, the critic, and the actor—that are trained concurrently from replayed experience without sharing gradients, as shown in Figure 3. To succeed across domains, these components need to accommodate different signal magnitudes and robustly balance terms in their objectives. This is challenging as we are not only targeting similar tasks within the same domain but aim to learn across different domains with fixed hyperparameters. This section first explains a simple transformation for predicting quantities of unknown orders of magnitude. We then introduce the world model, critic, and actor and their robust learning objectives. Specifically, we find that combining KL balancing and free bits enables the world model to learn without tuning, and scaling down large returns without amplifying small returns allows a fixed policy entropy regularizer. The differences to DreamerV2 are detailed in Appendix C.

Reconstructing inputs and predicting rewards and values can be challenging because their scale can vary across domains. Predicting large targets using a squared loss can lead to divergence whereas absolute and Huber losses 18 stagnate learning. On the other hand, normalizing targets based on running statistics 19 introduces non-stationarity into the optimization. We suggest symlog predictions as a simple solution to this dilemma. For this, a neural network f(x,θ)f(x,\theta) with inputs xx and parameters θ\theta learns to predict a transformed version of its targets yy. To read out predictions y^\hat{y} of the network, we apply the inverse transformation:

As shown in Figure 4, using the logarithm as transformation would not allow us to predict targets that take on negative values. Therefore, we choose a function from the bi-symmetric logarithmic family 20 that we name symlog as the transformation with the symexp function as its inverse:

The symlog function compresses the magnitudes of both large positive and negative values. Unlike the logarithm, it is symmetric around the origin while preserving the input sign. This allows the optimization process to quickly move the network predictions to large values when needed. Symlog approximates the identity around the origin so that it does not affect learning of targets that are already small enough. For critic learning, a more involved transformation has previously been proposed 21, which we found less effective on average across domains.

DreamerV3 uses symlog predictions in the decoder, the reward predictor, and the critic. It also squashes inputs to the encoder using the symlog function. Despite its simplicity, this approach robustly and quickly learns across a diverse range of environments. With symlog predictions, there is no need for truncating large rewards 18, introducing non-stationary through reward normalization 19, or adjusting network weights when new extreme values are detected 22.

World Model Learning

The world model learns compact representations of sensory inputs through autoencoding 23, 24 and enables planning by predicting future representations and rewards for potential actions. We implement the world model as a Recurrent State-Space Model (RSSM) 14, as shown in Figure 3. First, an encoder maps sensory inputs xtx_{t} to stochastic representations ztz_{t}. Then, a sequence model with recurrent state hth_{t} predicts the sequence of these representations given past actions at−1a_{t-1}. The concatenation of hth_{t} and ztz_{t} forms the model state from which we predict rewards rtr_{t} and episode continuation flags ct∈{0,1}c_{t}\in\{0,1\} and reconstruct the inputs to ensure informative representations:

The prediction loss trains the decoder and reward predictor via the symlog loss and the continue predictor via binary classification loss. The dynamics loss trains the sequence model to predict the next representation by minimizing the KL divergence between the predictor pϕ(zt  ∣  ht)p_{\phi}(z_{t}\;|\;h_{t}) and the next stochastic representation qϕ(zt  ∣  ht,xt)q_{\phi}(z_{t}\;|\;h_{t},x_{t}). The representation loss trains the representations to become more predictable if the dynamics cannot predict their distribution, allowing us to use a factorized dynamics predictor for fast sampling when training the actor critic. The two losses differ in the stop-gradient operator sg⁡(⋅)\operatorname{sg}(\cdot) and their loss scale. To avoid a degenerate solution where the dynamics are trivial to predict but contain not enough information about the inputs, we employ free bits28 by clipping the dynamics and representation losses below the value of 1 nat ≈\approx 1.44 bits. This disables them while they are already minimized well to focus the world model on its prediction loss:

Previous world models require scaling the representation loss differently based on the visual complexity of the environment. Complex 3D environments contain details unnecessary for control and thus prompt a stronger regularizer to simplify the representations and make them more predictable. In 2D games, the background is often static and individual pixels may matter for the task, requiring a weak regularizer to perceive fine details. We find that combining free bits with a small scale for the representation loss resolve this dilemma, allowing for fixed hyperparameters across domains. Moreover, symlog predictions for the decoder unify the gradient scale of the prediction loss across environments, further stabilizing the trade-off with the representation loss.

We occasionally observed spikes the in KL losses in earlier experiments, consistent with reports for deep variational autoencoders 29, 30. To prevent this, we parameterize the categorical distributions of the encoder and dynamics predictor as mixtures of 1% uniform and 99% neural network output, making it impossible for them to become near deterministic and thus ensuring well-scaled KL losses. Further model details and hyperparameters are summarized in Appendix W.

Actor Critic Learning

The actor and critic neural networks learn behaviors purely from abstract sequences predicted by the world model 12, 13. During environment interaction, we select actions by sampling from the actor network without lookahead planning. The actor and critic operate on model states st≐{ht,zt}s_{t}\doteq\{h_{t},z_{t}\} and thus benefit from the Markovian representations learned by the world model. The actor aims to maximize the expected return Rt≐∑τ=0∞γτrt+τR_{t}\doteq\textstyle\sum_{\tau=0}^{\infty}\gamma^{\tau}r_{t+\tau} with a discount factor γ=0.997\gamma=0.997 for each model state. To consider rewards beyond the prediction horizon T=16T=16, the critic learns to predict the return of each state under the current actor behavior:

Starting from representations of replayed inputs, the dynamics predictor and actor produce a sequence of imagined model states s1:Ts_{1:T}, actions a1:Ta_{1:T}, rewards r1:Tr_{1:T}, and continuation flags c1:Tc_{1:T}. To estimate returns that consider rewards beyond the prediction horizon, we compute bootstrapped λ\lambda-returns that integrate the predicted rewards and values 31, 32:

A simple choice for the critic loss function would be to regress the λ\lambda-returns via squared error or symlog predictions. However, the critic predicts the expected value of a potentially widespread return distribution, which can slow down learning. We choose a discrete regression approach for learning the critic based on twohot encoded targets 33, 34, 35, 36 that let the critic maintain and refine a distribution over potential returns. For this, we transform returns using the symlog function and discretize the resulting range into a sequence BB of K=255K=255 equally spaced buckets bib_{i}. The critic network outputs a softmax distribution pψ(bi  ∣  st)p_{\psi}(b_{i}\;|\;s_{t}) over the buckets and its output is formed as the expected bucket value under this distribution. Importantly, the critic can predict any continuous value in the interval because its expected bucket value can fall between the buckets:

To train the critic, we symlog transform the targets RtλR^{\lambda}_{t} and then twohot encode them into a soft label for the softmax distribution produced by the critic. Twohot encoding is a generalization of onehot encoding to continuous values. It produces a vector of length ∣B∣|B| where all elements are except for the two entries closest to the encoded continuous number, at positions kk and k+1k+1. These two entries sum up to 11, with more weight given to the entry that is closer to the encoded number:

Given twohot encoded targets yt=sg⁡(twohot⁡(symlog⁡(Rtλ)))y_{t}=\operatorname{sg}(\operatorname{twohot}(\operatorname{symlog}(R^{\lambda}_{t}))), where sg⁡(⋅)\operatorname{sg}(\cdot) stops the gradient, the critic minimizes the categorical cross entropy loss for classification with soft targets:

We found discrete regression for the critic to accelerate learning especially in environments with sparse rewards, likely because of their bimodal reward and return distributions. We use the same discrete regression approach for the reward predictor of the world model.

Because the critic regresses targets that depend on its own predictions, we stabilize learning by regularizing the critic towards predicting the outputs of an exponentially moving average of its own parameters. This is similar to target networks used previously in reinforcement learning 18 but allows us to compute returns using the current critic network. We further noticed that the randomly initialized reward predictor and critic networks at the start of training can result in large predicted rewards that can delay the onset of learning. We initialize the output weights of the reward predictor and critic to zeros, which effectively alleviates the problem and accelerates early learning.

Actor Learning

The actor network learns to choose actions that maximize returns while ensuring sufficient exploration through an entropy regularizer 37. However, the scale of this regularizer heavily depends on the scale and frequency of rewards in the environment, which has been a challenge for previous algorithms 38. Ideally, we would like the policy to explore quickly in the absence of nearby returns without sacrificing final performance under dense returns.

To stabilize the scale of returns, we normalize them using moving statistics. For tasks with dense rewards, one can simply divide returns by their standard deviation, similar to previous work 19. However, when rewards are sparse, the return standard deviation is generally small and this approach would amplify the noise contained in near-zero returns, resulting in an overly deterministic policy that fails to explore. Therefore, we propose propose to scale down large returns without scaling up small returns. We implement this idea by dividing returns by their scale SS, for which we discuss multiple choices below, but only if they exceed a minimum threshold of 11. This simple change is the key to allowing a single entropy scale η=3⋅10−4\eta=3\cdot 10^{-4} across dense and sparse rewards:

We follow DreamerV2 27 in estimating the gradient of the first term by stochastic backpropagation for continuous actions and by reinforce 39 for discrete actions. The gradient of the second term is computed in closed form.

In deterministic environments, we find that normalizing returns by their exponentially decaying standard deviation suffices. However, for heavily randomized environments, the return distribution can be highly non-Gaussian and contain outliers of large returns caused by a small number of particularly easy episodes, leading to an overly deterministic policy that struggles to sufficiently explore. To normalize returns while being robust to such outliers, we scale returns by an exponentially decaying average of the range from their 5th to their 95th batch percentile:

Because a constant return offset does not affect the objective, this is equivalent to an affine transformation that maps these percentiles to and 11, respectively. Compared to advantage normalization 19, scaling returns down accelerates exploration under sparse rewards without sacrificing final performance under dense rewards, while using a fixed entropy scale.

Results

We perform an extensive empirical study to evaluate the generality and scalability of DreamerV3 across diverse domains—with over 150 tasks—under fixed hyperparameters. We designed the experiments to compare DreamerV3 to the best methods in the literature, which are often specifically designed to the benchmark at hand. Moreover, we apply DreamerV3 to the challenging video game Minecraft. Appendix A gives an overview of the domains. For DreamerV3, we directly report the performance of the stochastic training policy and avoid separate evaluation runs using the deterministic policy, simplifying the setup. All DreamerV3 agents are trained on one Nvidia V100 GPU each, making the algorithm widely usable across research labs. The source code and numerical results are available on the project website: https://danijar.com/dreamerv3

To evaluate the generality of DreamerV3, we perform an extensive empirical evaluation across 7 domains that include continuous and discrete actions, visual and low-dimensional inputs, dense and sparse rewards, different reward scales, 2D and 3D worlds, and procedural generation. Figure 1 summarizes the results, with training curves and score tables included in the appendix. DreamerV3 achieves strong performance on all domains and outperforms all previous algorithms on 4 of them, while also using fixed hyperparameters across all benchmarks.

Proprio Control Suite This benchmark contains 18 continuous control tasks with low-dimensional inputs and a budget of 500K environment steps 40. The tasks range from classical control over locomotion to robot manipulation tasks. DreamerV3 sets a new state-of-the-art on this benchmark, outperforming D4PG 41, DMPO 42, and MPO 43.

Visual Control Suite This benchmark consists of 20 continuous control tasks where the agent receives only high-dimensional images as inputs and a budget of 1M environment steps 40, 27. DreamerV3 establishes a new state-of-the-art on this benchmark, outperforming DrQ-v2 44 and CURL 45 which additionally require data augmentations.

Atari 100k This benchmark includes 26 Atari games and a budget of only 400K environment steps, amounting to 100K steps after action repeat or 2 hours of real time 46. EfficientZero 47 holds the state-of-the-art on this benchmark by combining online tree search, prioritized replay, hyperparameter scheduling, and allowing early resets of the games; see Appendix T for an overview. Without this complexity, DreamerV3 outperforms the remaining previous methods such as the transformer-based IRIS 48, the model-free SPR 49, and SimPLe 46.

Atari 200M This popular benchmark includes 55 Atari video games with simple graphics and a budget of 200M environment steps 50. We use the sticky action setting 51. DreamerV3 outperforms DreamerV2 with a median score of 302% compared to 219%, as well as the top model-free algorithms Rainbow 52 and IQN 53 that were specifically designed for the Atari benchmark.

BSuite This benchmark includes 23 environments with a total of 468 configurations that are designed to test credit assignment, robustness to reward scale and stochasticity, memory, generalization, and exploration 54. DreamerV3 establishes a new state-of-the-art on this benchmark, outperforming Bootstrap DQN 55 as well as Muesli 56 with comparable amount of training. DreamerV3 improves over previous algorithms the most in the credit assignment category.

Crafter This procedurally generated survival environment with top-down graphics and discrete actions is designed to evaluate a broad range of agent abilities, including wide and deep exploration, long-term reasoning and credit assignment, and generalization 57. DreamerV3 sets a new state-of-the-art on this benchmark, outperforming PPO with the LSTM-SPCNN architecture 58, the object-centric OC-SA 58, DreamerV2 27, and Rainbow 52.

DMLab This domain contains 3D environments that require spatial and temporal reasoning 59. On 8 challenging tasks, DreamerV3 matches and exceeds the final performance on the scalable IMPALA agent 60 in only 50M compared to 10B environment steps, amounting to a data-efficiency gain of over 13000%. We note that IMPALA was not designed for data-efficiency but serves as a valuable baseline for the performance achievable by scalable RL algorithms without data constraints.

Scaling properties

Solving challenging tasks out of the box not only requires an algorithm that succeeds without adjusting hyperparameters, but also the ability to leverage large models to solve hard tasks. To investigate the scaling properties of DreamerV3, we train 5 model sizes ranging from 8M to 200M parameters. As shown in Figure 6, we discover favorable scaling properties where increasing the model size directly translates to both higher final performance and data-efficiency. Increasing the number of gradient steps further reduces the number of interactions needed to learn successful behaviors. These insights serve as practical guidance for applying DreamerV3 to new tasks and demonstrate the robustness and scalability of the algorithm.

Minecraft

Collecting diamonds in the open-world game Minecraft has been a long-standing challenge in artificial intelligence. Every episode in this game is set in a different procedurally generated 3D world, where the player needs to discover a sequence of 12 milestones with sparse rewards by foraging for resources and using them to craft tools. The environment is detailed in Appendix F. We following prior work 17 and increase the speed at which blocks break because a stochastic policy is unlikely to sample the same action often enough in a row to break blocks without regressing its progress by sampling a different action.

Because of the training time in this complex domain, tuning algorithms specifically for Minecraft would be difficult. Instead, we apply DreamerV3 out of the box with its default hyperparameters. As shown in Figure 1, DreamerV3 is the first algorithm to collect diamonds in Minecraft from scratch without using human data that was required by VPT 16. Across 40 seeds trained for 100M environment steps, DreamerV3 collects diamonds in 50 episode. It collects the first diamond after 29M steps and the frequency increases as training progresses. A total of 24 of the 40 seeds collect at least one diamond and the most successful agent collects diamonds in 6 episodes. The success rates for all 12 milestones are shown in Figure G.1.

Previous Work

Developing general-purpose algorithms has long been a goal of reinforcement learning research. PPO 19 is one of the most widely used algorithms and requires relatively little tuning but uses large amounts of experience due to its on-policy nature. SAC 38 is a popular choice for continuous control and leverages experience replay for higher data-efficiency, but in practice requires tuning, especially for its entropy scale, and struggles with high-dimensional inputs 61. MuZero 34 plans using a value prediction model and has achieved high performance at the cost of complex algorithmic components, such as MCTS with UCB exploration and prioritized replay. Gato 62 fits one large model to expert demonstrations of multiple tasks, but is only applicable to tasks where expert data is available. In comparison, we show that DreamerV3 masters a diverse range of environments trained with fixed hyperparameters and from scratch.

Minecraft has been a focus of recent reinforcement learning research. With MALMO 63, Microsoft released a free version of the popular game for research purposes. MineRL 15 offers several competition environments, which we rely on as the basis for our experiments. MineDojo 64 provides a large catalog of tasks with sparse rewards and language descriptions. The yearly MineRL competition supports agents in exploring and learning meaningful skills through a diverse human dataset 15. VPT 16 trained an agent to play Minecraft through behavioral cloning of expert data collected by contractors and finetuning using reinforcement learning, resulting in a 2.5% success rate of diamonds using 720 V100 GPUs for 9 days. In comparison, DreamerV3 learns to collect diamonds in 17 GPU days from sparse rewards and without human data.

Conclusion

This paper presents DreamerV3, a general and scalable reinforcement learning algorithm that masters a wide range of domains with fixed hyperparameters. To achieve this, we systematically address varying signal magnitudes and instabilities in all of its components. DreamerV3 succeeds across 7 benchmarks and establishes a new state-of-the-art on continuous control from states and images, on BSuite, and on Crafter. Moreover, DreamerV3 learns successfully in 3D environments that require spatial and temporal reasoning, outperforming IMPALA in DMLab tasks using 130 times fewer interactions and being the first algorithm to obtain diamonds in Minecraft end-to-end from sparse rewards. Finally, we demonstrate that the final performance and data-efficiency of DreamerV3 improve monotonically as a function of model size.

Limitations of our work include that DreamerV3 only learns to sometimes collect diamonds in Minecraft within 100M environment steps, rather than during every episode. Despite some procedurally generated worlds being more difficult than others, human experts can typically collect diamonds in all scenarios. Moreover, we increase the speed at which blocks break to allow learning Minecraft with a stochastic policy, which could be addressed through inductive biases in prior work. To show how far the scaling properties of DreamerV3 extrapolate, future implementations at larger scale are necessary. In this work, we trained separate agents for all tasks. World models carry the potential for substantial transfer between tasks. Therefore, we see training larger models to solve multiple tasks across overlapping domains as a promising direction for future investigations.

We thank Oleh Rybkin, Mohammad Norouzi, Abbas Abdolmaleki, John Schulman, and Adam Kosiorek for insightful discussions. We thank Bobak Shahriari for training curves of baselines for proprioceptive control, Denis Yarats for training curves for visual control, Surya Bhupatiraju for Muesli results on BSuite, and Hubert Soyer for providing training curves of the original IMPALA experiments. We thank Daniel Furrer, Andrew Chen, and Dakshesh Garambha for help with using Google Cloud infrastructure for running the Minecraft experiments.

References

Appendices

psection[2.3em]\contentslabel2.3em\contentspage \startcontents[sections] \printcontents[sections]p1

Appendix A Benchmark Overview

colspec = | L6em | C3.5em C3.5em C3.5em C3.5em C3.5em C3.5em C3.5em |, row1 = font=,

Benchmark Tasks Env Steps Action Repeat Env Instances Train Ratio GPU Days Model Size DMC Proprio 18 500K 2 4 512 0<1 S DMC Vision 20 1M 2 4 512 0<1 S Crafter 1 1M 1 1 512 02 XL BSuite 23 — 1 1 1024 0<1 XL Atari 100K 26 400K 4 1 1024 0<1 S Atari 200M 55 200M 4 8 64 16 XL DMLab 8 50M 4 8 64 04 XL Minecraft 1 100M 1 16 16 17 XL

colspec = | L10em | C4.5em C4.5em C4.5em C4.5em C4.5em |, row1 = font=,

Dimension XS S M L XL GRU recurrent units 256 512 1024 2048 4096 CNN multiplier 24 32 48 64 96 Dense hidden units 256 512 640 768 1024 MLP layers 1 2 3 4 5 Parameters 8M 18M 37M 77M 200M

DreamerV3 builds upon the DreamerV2 algorithm 27. This section describes the main changes that we applied in order to master a wide range of domains with fixed hyperparameters and enable robust learning on unseen domains.

Symlog predictions We symlog encode inputs to the world model and use symlog predictions with squared error for reconstructing inputs. The reward predictor and critic use twohot symlog predictions, a simple form of distributional reinforcement learning 33.

World model regularizer We experimented with different approaches for removing the need to tune the KL regularizer, including targeting a fixed KL value. A simple yet effective solution turned out to combine KL balancing that was introduced in DreamerV2 27 with free bits 28 that were used in the original Dreamer algorithm 65. GECO 66 was not useful in our case because what constitutes “good” reconstruction error varies widely across domains.

Policy regularizer Using a fixed entropy regularizer for the actor was challenging when targeting both dense and sparse rewards. Scaling large return ranges down to the $$ interval, without amplifying near-zero returns, overcame this challenge. Using percentiles to ignore outliers in the return range further helped, especially for stochastic environments. We did not experience improvements from regularizing the policy towards its own EMA 67 or the CMPO regularizer 56.

Unimix categoricals We parameterize the categorical distributions for the world model representations and dynamics, as well as for the actor network, as mixtures of 1% uniform and 99% neural network output 68, 69 to ensure a minimal amount of probability mass on every class and thus keep log probabilities and KL divergences well behaved.

Architecture We use a similar network architecture but employ layer normalization 70 and SiLU 71 as the activation function. For better framework support, we use same-padded convolutions with stride 2 and kernel size 3 instead of valid-padded convolutions with larger kernels. The robustness of DreamerV3 allowed us to use large networks that contributed to its performance.

Critic EMA regularizer We compute λ\lambda-returns using the fast critic network and regularize the critic outputs towards those of its own weight EMA instead of computing returns using the slow critic. However, both approaches perform similarly in practice.

Replay buffer DreamerV2 used a replay buffer that only replays time steps from completed episodes. To shorten the feedback loop, DreamerV3 uniformly samples from all inserted subsequences of size batch length regardless of episode boundaries.

Hyperparameters The hyperparameters of DreamerV3 were tuned to perform well across both the visual control suite and Atari 200M at the same time. We verified their generality by training on new domains without further adjustment, including Crafter, BSuite, and Minecraft.

We also experimented with constrained optimization for the world model and policy objectives, where we set a target value that the regularizer should take on, on average across states. We found combining this approach with limits on the allowed regularizer scales to perform well for the world model, at the cost of additional complexity. For the actor, choosing a target randomness of 40%40\%—where 0%0\% corresponds to the most deterministic and 100%100\% to the most random policy—learns robustly across domains but prevents the policy from converging to top scores in tasks that require speed or precision and slows down exploration under sparse rewards. The solutions in DreamerV3 do not have these issues intrinsic to constrained optimization formulations.

Appendix D Ablation Curves

Use KL balancing but no free bits, equivalent to setting the constants in Equation 4 from 1 to 0. This objective was used in DreamerV2 27.

NoKLBalance

NoObsSymlog

This ablation removes the symlog encoding of inputs to the world model and also changes the symlog MSE loss in the decoder to a simple MSE loss. Because symlog encoding is only used for vector observations, this ablation is equivalent to DreamerV3 on purely image-based environments.

TargetKL

Critic Ablations

Instead of normalizing rewards, normalize rewards by dividing them by a running standard deviation and clipping them beyond a magnitude of 10.

ContRegression

Using MSE symlog predictions for the reward and value heads.

SqrtTransform

Using two-hot discrete regression with the asymmetric square root transformation introduced by R2D2 21 and used in MuZero 34.

SlowTarget

Instead of using the fast critic for computing returns and training it towards the slow critic, use the slow critic for computing returns 18.

Actor Ablations

Normalize returns directly based on the range between percentiles 5 to 95 with a small epsilon in the denominator, instead of by the maximum of 1 and the percentile range. This way, not only large returns are scaled down but also small returns are scaled up.

AdvantageStd

Advantage normalization as commonly used, for example in PPO 19 and Muesli 56. However, scaling advantages without also scaling the entropy regularizer changes the trade-off between return and entropy in a way that depends on the scale of advantages, which in turn depends on how well the critic currently predicts the returns.

ReturnStd

Instead of normalizing returns by the range between percentiles 5 to 95, normalize them by their standard deviation. When rewards are large but sparse, the standard deviation is small, scaling up the few large returns even further.

TargetEntropy

Target a policy randomness of 40% on average across imagined states by increasing or decreasing the entropy scale η\eta by 10% when the batch average of the randomness falls below or exceeds the tolerance of 10% around the target value. The entropy scale is limited to the range [10−3,3⋅10−2][10^{-3},3\cdot 10^{-2}]. Policy randomness is the policy entropy mapped to range from 0% (most deterministic allowed by action distribution parameterization) to 100% (most uniform). Multiplicatively, instead of additively, adjusting the regularizer strength allows the scale to quickly move across orders of magnitude, outperforming the target entropy approach of SAC 38 in practice. Moreover, targeting a randomness value rather than an entropy value allows sharing the hyperparameter across domains with discrete and continuous actions.

Appendix F Minecraft Environment

With 100M monthly active users, Minecraft is one of the most popular video games worldwide. Minecraft features a procedurally generated 3D world of different biomes, including plains, forests, jungles, mountains, deserts, taiga, snowy tundra, ice spikes, swamps, savannahs, badlands, beaches, stone shores, rivers, and oceans. The world consists of 1 meter sized blocks that the player and break and place. There are about 30 different creatures that the player can interact and fight with. From gathered resources, the player can use 379 recipes to craft new items and progress through the technology tree, all while ensuring safety and food supply to survive. There are many conceivable tasks in Minecraft and as a first step, the research community has focused on the salient task of obtaining a diamonds, a rare item found deep underground and requires progressing through the technology tree.

Environment

We built the Minecraft Diamond environment on top of MineRL to define a flat categorical action space and fix issues we discovered with the original environments via human play testing. For example, when breaking diamond ore, the item sometimes jumps into the inventory and sometimes needs to be collected from the ground. The original environment terminates episodes when breaking diamond ore so that many successful episodes end before collecting the item and thus without the reward. We remove this early termination condition and end episodes when the player dies or after 36000 steps, corresponding to 30 minutes at the control frequency of 20Hz. Another issue is that the jump action has to be held for longer than one control step to trigger a jump, which we solve by keeping the key pressed in the background for 200ms. We built the environment on top of MineRL v0.4.4 15, which offers abstract crafting actions. The Minecraft version is 1.11.2.

Rewards

We follow the same sparse reward structure of the MineRL competition environment that rewards 12 milestones leading up to the diamond, namely collecting the items log, plank, stick, crafting table, wooden pickaxe, cobblestone, stone pickaxe, iron ore, furnace, iron ingot, iron pickaxe, and diamond. The reward for each item is only given once per episode, and the agent has to learn autonomously that it needs to collect some of the items multiple times to achieve the next milestone. To make the return curves easy to interpret, we give a reward of +1+1 for each milestone instead of scaling rewards based on how valuable each item is. Additionally, we give a small reward of −0.01-0.01 for each lost heart and +0.01+0.01 for each restored heart, but we did not investigate whether this was helpful.

Inputs

The sensory inputs include the 64×64×364\times 64\times 3 RGB first-person camera image, the inventory counts as a vector with one entry for each of the game’s over 400 items, the vector of maximum inventory counts since episode begin to tell the agent which milestones it has already achieved, a one-hot vector indicating the equipped item, and scalar inputs for the health, hunger, and breath levels.

Actions

The MineRL environment provides a dictionary action space and delegates choosing a simple action space to the user. We use a flat categorical action space with 25 actions for walking in four directions, turning the camera in four directions, attacking, jumping, placing items, crafting items near a placed crafting table, smelting items near a placed furnace, and equipping crafted tools. Looking up and down is restricted to the range −60-60 to +60+60 degrees. The action space is specific to the diamond task and does not allow the agent to craft all of the 379 recipes. For multi-task learning, a larger factorized action space as available in MineDojo 64 would likely be beneficial.

Break Speed Multiplier

We follow Kanitscheider et al. 17 in allowing the agent to break blocks in fewer time steps. Breaking blocks in Minecraft requires holding the same key pressed for sometimes hundreds of time steps. Briefly switching to a different key will reset this progress. Without an inductive bias, stochastic policies will almost never sample the same action this often during exploration under this parameterization. To circumvent this issue, we set the break speed multiplier option of MineRL to 100. In the future, inductive biases such as learning action repeat as part of the agent 74 could overcome this caveat.

Appendix G Minecraft Item Rates

Across 40 seeds trained for 100M steps, DreamerV3 obtained the maximum episode score—that includes collecting at least one diamond—50 times. It achieves this score the first time at 29.3M steps and as expected the frequency increases over time. Diamonds can also rarely be found by breaking village chests, but those episodes do not achieve the maximum score and thus are not included in this statistic. A total of 24 out of 40 seeds achieve the maximum episode score at least once, and the most successful seed achieved the maximum score 6 times. Across all seeds, the median number of environment steps until collecting the first diamond is 74M, corresponding to 42 days of play time at 20 Hz.