Tianshou: a Highly Modularized Deep Reinforcement Learning Library
Jiayi Weng, Huayu Chen, Dong Yan, Kaichao You, Alexis Duburcq, Minghao Zhang, Yi Su, Hang Su, Jun Zhu
Introduction
Recent advances in deep reinforcement learning (DRL) have ignited enthusiasm of both academia and industry. This is accompanied by the flourish of many newly-emerged DRL algorithms (Mnih et al., 2015; Silver et al., 2017; Duan et al., 2016), together with numerous libraries that try to provide reference implementations with the representative ones including RLlib (Liang et al., 2018), rlpyt (Stooke and Abbeel, 2019), Stable-Baselines3 (Raffin et al., 2021), MushroomRL (D’Eramo et al., 2020), and PFRL (Fujita et al., 2021).
Though most DRL libraries are comprehensive, many researchers still tend to use their own DRL code base for fast prototyping in practice. The reasons behind this are various. Some libraries may choose to highly encapsulate the supported algorithms, leaving out several options to be tweaked. This assists algorithms’ application, but harms the flexibility because it is impossible to provide exhaustive options. The usability is also a concern. Several libraries prioritize supporting highly distributed training, but bring complex code structure and difficulties in debugging. However, parallelized data sampling is still needed in research because of the typically unbalanced numbers of CPUs and GPUs in a server. Another reason might be that many libraries are only comprehensive for a certain type of algorithms (e.g., only support online or offline algorithms) to keep a unified API.
To address these issues, we present Tianshou, a highly modularized Python library for deep reinforcement learning based on PyTorch. Tianshou has the following characteristics:
Highly modularized. Tianshou aims to provide building blocks rather than training scripts. It can be easily used for fast prototyping because the shared infrastructure commonly used in DRL is factored out (Figure 1). Users only need to change a few variables to apply techniques commonly used in DRL (e.g., parallel data sampling).
Reliable. Tianshou has a code coverage of 94%. Every commit to Tianshou will go through unit tests on multiple platforms. We have also released a systematic benchmark of Gym’s MuJoCo environmenthttps://tianshou.readthedocs.io/en/master/tutorials/benchmark.html (Todorov et al., 2012; Brockman et al., 2016). In this benchmark, Tianshou incorporates a comprehensive set of DRL techniques for 8 benchmarked algorithms and scores 15% higher on average in terms of median performance compared with reference implementations.
Comprehensive. Besides comprehensive model-free algorithms, Tianshou also supports offline learning and many other DRL techniques such as GAIL (Ho and Ermon, 2016) and ICM (Pathak et al., 2017). Moreover, through a unified Python interface, Tianshou formulates the data collecting (e.g., both synchronous and asynchronous environment execution) and agent training paradigms in DRL. Lastly, Tianshou has plentiful functionalities that may extend its application (see Section 2).
Architecture of Tianshou
In this section, we will briefly introduce Tianshou’s architecture as illustrated in Figure 1.
Standardization of the Training Process. We standardize the training paradigms of mainstream DRL algorithms by considering different experience replay mechanisms and classify them into three types: on-policy training, off-policy training, and offline learning. We use a replay buffer to store the transitions and a collector to collect transition data into the buffer. We use the policy’s update function to update the parameter (Figure 2).
Parallel Computing Infrastructure. In concurrent research, Stooke and Abbeel (2019) addresses two phases of parallelization in DRL: environment sampling and agent training. Tianshou targets small- to medium-scale research, so it focuses on the first one. Following Clemente et al. (2017), we adopt their parallelization technique to balance simulation and inference loads. Note that our contributions to parallel sampling schemes exceed this work by allowing asynchronized sampling as an alternative way to ease the straggler effect instead of only stacking environment instances per process. Thanks to the modularized design, Tianshou can easily support the C++-based vectorized environment EnvPool (Weng et al., 2022) with free speed up.
Utilities. Tianshou intends to relieve users from imperceptible details critical to a desirable performance and, hence, incorporates a comprehensive set of DRL techniques as its infrastructure. Techniques include partial-episode bootstrapping (Pardo et al., 2018), observation/value normalization (van Hasselt et al., 2016), automatic action scaling and GAE (Schulman et al., 2016), etc. Besides, Tianshou has plentiful extra functionalities that users might find helpful. For instance, Tianshou has customizable loggers compatible with TensorBoard and W&B. Recurrent state representation, prioritized experience replay, training resumption, and buffer serialization are also supported.
Reproduction Scripts and Performance. We have released Tianshou’s OpenAI Gym MuJoCo task suite benchmark, covering 8 classic algorithms and 9 environments. In this benchmark, Tianshou scores 15% higher on average compared with multiple reference implementations in terms of 9 environments’ median performance, demonstrating its reliability. For discrete action space problems, we also provide example code and results with 7 supported algorithms in 7 Atari environments. All experiments are done with 10 random seeds. Some libraries (Fujita et al., 2021) devote themselves to faithfully replicating existing papers, while Tianshou aims to present an as-consistent-as-possible set of hyperparameters and low-level designs. While leaving the core algorithm untouched, we try to incorporate several known tricks in a specific algorithm to all similar algorithms supported by Tianshou. Hopefully, this will facilitate comparisons between algorithms.
Usability
Tianshou is lightweight and easy to install. Users can simply install Tianshou via Pip or Conda on different platforms (Windows, macOS, Linux). Full API documentation and a series of tutorials are provided at https://tianshou.readthedocs.io/. Only a few lines of code are required to start a simple experiment. Tianshou also strictly follows the PEP8 code style with the code commented and data type annotated. Contributing guidelines and extensive unit tests with GitHub Actions, including code-style, type, and performance checks, help Tianshou maintain its code quality.
Comparison to Related Works
While several TensorFlow-based DRL libraries (Kuhnle et al., 2017; Plappert, 2016; Caspi et al., 2017) are also worth mentioning, we limit our comparison to a few DRL libraries with the PyTorch backend due to page limit. RLlib (Liang et al., 2018) and rlpyt (Stooke and Abbeel, 2019) are libraries designed to be high-throughput software and support both multi-CPU parallel sampling and multi-GPU optimization, while Stable-Baselines3 (Raffin et al., 2021), PFRL (Fujita et al., 2021) and Tianshou focus on small- to medium-scale application of DRL algorithms and support only parallel sampling. Libraries like MushroomRL (D’Eramo et al., 2020) are intentionally designed to be research-friendly, so no parallelization is supported. This leads to very different design choices and results in different code complexity. In terms of supported algorithms, most libraries are comprehensive, but each has a different focus. MushroomRL supports both classic RL and DRL algorithms to facilitate research. d3rlpy (Takuma Seno, 2021) prioritizes supporting offline DRL algorithms and other libraries mainly support online algorithms. While many libraries above are modular, Tianshou achieves modularity mainly by factoring out the infrastructure in DRL, compared to several libraries that focus on offering one highly encapsulated API for each algorithm (Raffin et al., 2021; Takuma Seno, 2021). Above all, the architecture of PFRL is most similar to Tianshou. However, they still have key differences like the implementation of lower-level data container Batch and how to support sequence Buffers for RNN. Other engineering features are compared in Tianshou’s GitHub repositoryhttps://github.com/thu-ml/tianshou/blob/master/README.md#why-tianshou.
Conclusion
This paper briefly describes Tianshou, a flexible and reliable implementation of a modular DRL library. Tianshou sets up a framework for DRL research by factoring out the shared infrastructure commonly used in DRL as building blocks. We have also released a MuJoCo benchmark, covering many classic algorithms, demonstrating Tianshou’s reliability.
We thank Haosheng Zou for his early work on TensorFlow-based Tianshou before version 0.1.1. We thank Peng Zhong, Qiang He, Chengqi Duan, Qing Xiao, Qifan Li, Yan Li, and others for their valuable contributions to Tianshou.
This work was supported by the National Key Research and Development Program of China ( 2020AAA0106000, 2020AAA0104304, 2020AAA0106302, 2021YFB2701000), NSFC Projects (Nos. 62061136001, 62076147, U19B2034, U1811461, U19A2081, 61972224), Beijing NSF Project (No. JQ19016), BNRist (BNR2022RC01006), Tsinghua Institute for Guo Qiang, Beijing Academy of Artificial Intelligence (BAAI), Tsinghua-Huawei Joint Research Program, and the High Performance Computing Center, Tsinghua University.