Learning to Paint With Model-based Deep Reinforcement Learning

Zhewei Huang, Wen Heng, Shuchang Zhou

Introduction

Painting, being an important form of visual art, symbolizes human wisdom and creativity. In recent centuries, artists have used a diverse array of tools to create their masterpieces. But it’s hard for people to master this skill without spending a large amount of time in proper training. Therefore, teaching machines to paint is a challenging task and helps to shed light on the mystery of painting. Furthermore, the study of this topic can help us build painting assistant tools.

We train an artificial intelligence painting agent that can paint strokes on a canvas in sequence to generate a painting that resembles a given image. Neural networks are used to produce parameters that control the position, shape, color, and transparency of strokes. Previous works have studied teaching machines to learn painting-related skills, such as sketching , doodling and writing characters . In contrast, we aim to teach machines to handle more complex tasks, such as painting portraits of humans and natural scenes in the real world, which have rich textures and complex structural compositions.

We address three challenges for training an agent to paint real-world images. First, painting like humans requires an agent to have the ability to decompose a given target image into an ordered sequence of strokes. The agent needs to parse the target image visually, understand the current status of the canvas, and have foresightful plans about future strokes. To achieve this planning, one method is to give the supervised loss for stroke decomposition at each step, as in . However, such a method require ground truth stroke decomposition, which is hard to define. Also, texture-rich image painting usually requires hundreds of strokes to generate a painting that resembles a target image, which is tens of times more than doodling, sketching or character writing require and increases the difficulty of planning. To handle the ill-definedness of the problem, and the long-term planning challenge, we propose using reinforcement learning (RL) to train the agent, because RL can maximize the cumulative rewards of a whole painting process rather than minimizing supervised loss at each step. Experiments show that an RL agent can build plans for stroke decomposition with hundreds of steps. Moreover, we apply the adversarial training strategy to improve the pixel-level quality of the generated images, as the strategy has proved effective in other image generation tasks .

Second, we design continuous stroke parameter space, including stroke location, color, and transparency, to improve the painting quality. Previous works design discrete stroke parameter spaces and each parameter has only a limited number of choices, which fall short for texture-rich paintings. Instead, we adopt the Deep Deterministic Policy Gradient (DDPG) which copes well with the continuous action space of the agent.

Third, we build an efficient differentiable neural renderer that can simulate painting of hundreds of strokes on the canvas. Most previous works paint by interacting with undifferentiable painting simulation environments, which are good as renders but fail to provide detailed feedback about the generated images. Instead, we train a neural network that directly maps stroke parameters to stroke paintings. The renderer can also be adapted to different stroke designs like triangle and circles by changing the generation patterns. Moreover, the differential renderer can be combined with DDPG into a single model-based DRL that can be trained in an end-to-end fashion, which significantly boosts both the painting quality and convergence speed.

In summary, our contributions are three-fold:

We approach the painting task with the model-based DRL algorithm and build agents that decompose the target image into hundreds of strokes in sequence which can recreate a painting on canvas.

We build differentiable neural renderers for efficient painting and flexible support of different stroke designs, e.g. Bézier curve, triangle, and circle. The neural renderer contributes to the painting quality by allowing training model-based DRL agent in an end-to-end fashion.

Experiments show that the proposed painting agent can handle multiple types of target images well, including handwritten digits, streetview house numbers, human portraits, and natural scene images.

Related work

Stroke-based rendering (SBR) is a method of non-photorealistic imagery that recreates images by placing discrete drawing elements such as paint strokes or stipples on canvas. Most SBR algorithms solve the stroke decomposition problem by Greedy Search on every single step or require user interaction. Haeberli et al. propose a semiautomatic method which requires the user to set parameters to control the shape of the strokes and select the positions for each stroke. Litwinowicz et al. propose a single-layer painter-like rendering which places the brush strokes on a grid in the image plane, with randomly perturbed positions. Some work also studies the effects of using different stroke designs and the related problem of generating animations from video .

Recent works use RL to improve the stroke decomposition of images. SPIRAL is an adversarially trained DRL agent that learn structures in images, but fails to recover the details of human portraits. StrokeNet combines differentiable renderer and recurrent neural network (RNN) to train agents to paint but fails to generalize on color images. Doodle-SDQ trains the agents to emulate human doodling with DQN. Earlier, Sketch-RNN uses sequential datasets to achieve good results in sketch drawings. Artist Agent explores using RL for the automatic generation of a single brush stroke.

Painting Agent

The goal of the painting agent is decomposing the given target image into strokes that can recreate the image on the canvas. To imitate the painting process of humans, the agent is designed to predict the next stroke based on observing the current state of the canvas and the target image. However, the stroke at each step needs to be well compatible with previous strokes and future strokes to reduce the number of strokes for finishing the painting. We postulate that the agent should maximize the cumulative rewards after finishing the given number of strokes, rather than the gain of current stroke. To achieve this delayed-reward design, we employ a DRL framework, with the diagrams for the overall architecture shown in Figure 2.

In the framework, we model the painting process as a sequential decision-making task, which is described in Section 3.2. And to build the feedback mechanism, we use a neural renderer to help generate detailed rewards for training the agent, which is described in Section 3.3.

2 The Model

Given a target image II and an empty canvas C0C_{0}, the agent aims to find a stroke sequence (a0,a1,...,an−1)(a_{0},a_{1},...,a_{n-1}), where rendering ata_{t} on CtC_{t} can get Ct+1C_{t+1}. After rendering these strokes in sequence, we get the final painting CnC_{n}, which should be visually similar to II as much as possible. We model this task as a Markov Decision Process with a state space S\mathcal{S}, an action space A\mathcal{A}, a transition function trans⁡(st,at)\operatorname{trans}(s_{t},a_{t}) and a reward function r(st,at)r(s_{t},a_{t}). The details of these components are specified next.

State and Transition Function The state space is constructed by all possible information that the agent can observe in the environment. We separate a state into three parts: states of the canvas, the target image, and the step number. Formally, st=(Ct,I,t)s_{t}=(C_{t},I,t). CtC_{t} and II are bitmaps and the step number tt acts as additional information to instruct the agent the remaining number of steps. The transition function , st+1=trans⁡(st,at)s_{t+1}=\operatorname{trans}(s_{t},a_{t}) gives the transition process between states, which is implemented by painting a stroke on the current canvas.

Action An action ata_{t} of the painting agent is a set of parameters that control the position, shape, color and transparency of the stroke that would be painted at step tt. We define the behavior of an agent as a policy function π\pi that maps states to deterministic actions, i.e. π ⁣:S→A\pi\colon\mathcal{S}\to\mathcal{A}. At step tt, the agent observes state sts_{t} before predicting the parameters of the next stroke ata_{t}. The state evolves based on the transition function st+1=trans⁡(st,at)s_{t+1}=\operatorname{trans}(s_{t},a_{t}), which runs for nn steps.

Reward Selecting a suitable metric to measure the difference between the current canvas and the target image is found to be crucial for training a painting agent. The reward is designed as follows,

where r(st,at)r(s_{t},a_{t}) is the reward at step tt, LtL_{t} is the measured loss between II and the CtC_{t} and Lt+1L_{t+1} is the measured loss between II and the Ct+1C_{t+1}. In this work, LL is formulated as the discriminator score that is defined in Section 3.3.3.

To make the final canvas resemble the target image, the agent should be driven to maximize the cumulative rewards in the whole episode. At each step, the objective of the agent is to maximize the sum of discounted future reward Rt=∑i=tTγ(i−t)r(si,ai)R_{t}=\sum_{i=t}^{T}\gamma^{(i-t)}r(s_{i},a_{i}) with a discounting factor γ∈\gamma\in.

3 Learning

In this section, we introduce how to train the agent using the model-based DDPG algorithm.

We first describe the original DDPG, then introduce building model-based DDPG for efficient agent training.

As we use continuous parameters for strokes, the action space in the painting task is continuous and of high dimensions. Discretizing the action space to adapt some DRL methods, such as DQN and PG, will lose the precision of stroke representation and require many efforts in manual structure design to cope with the explosion of parameter combinations in discrete space. In contrast, DPG uses deterministic policy to resolve the difficulties caused by high-dimensional continuous action space, and DDPG is its variant using Neural Networks.

In the original DDPG, there are two networks: the actor π(s)\pi(s) and critic Q(s,a)Q(s,a). The actor models a policy π\pi that maps a state sts_{t} to action ata_{t}. The critic estimates the expected reward for the agent taking action ata_{t} at state sts_{t}, which is trained using Bellman equation (2) as in Q-learning and the data is sampled from an experience replay buffer:

Here r(st,at)r(s_{t},a_{t}) is a reward given by the environment when performing action ata_{t} at state sts_{t}. The actor π(st)\pi(s_{t}) is trained to maximize the critic’s estimated Q(st,π(st))Q(s_{t},\pi(s_{t})). In other words, the actor decides a stroke for each state. Based on the current canvas and the target image, the critic predicts an expected reward for the stroke. The critic is optimized to estimate more accurate expected rewards.

We cannot train a good-performance painting agent using original DDPG because it’s hard for the agent to model the complex environment well that is composed of any types of real-world images during learning. The World Model is a method to make agent understand the environments effectively. Similarly, we design a neural renderer so that the agent can observe a modeled environment. Then it can explore the environment and improve its policy efficiently. We term the DDPG with the actor that can get access to the gradients from environments as model-based DDPG. The difference between the two algorithms is visually shown in Figure 4.

The optimization of the agent using the model-based DDPG is different from that using the original DDPG. At step tt, the critic takes st+1s_{t+1} as input rather than both of sts_{t} and ata_{t}. The critic still predicts the expected reward for the state but no longer includes the reward caused by the current action. The new expected reward is a value function V(st)V(s_{t}) trained using discounted reward:

Here r(st,at)r(s_{t},a_{t}) is the reward when performing action ata_{t} based on sts_{t}. The actor π(st)\pi(s_{t}) is trained to maximize r(st,π(st))+V(trans⁡(st,π(st)))r(s_{t},\pi(s_{t}))+V(\operatorname{trans}(s_{t},\pi(s_{t}))). The transition function st+1=trans⁡(st,at)s_{t+1}=\operatorname{trans}(s_{t},a_{t}) is the differentiable renderer.

3.2 Action Bundle

Frame Skip is a powerful trick for many RL tasks, by restricting the agent to only observe the environment and acts once every kk frames rather than one frame. The trick makes the agents have a better ability to learn associations between more temporally distant states and actions. The agent predicts one action and reuse it at the next k−1k-1 frames instead and achieves better performance with less computation cost.

Inspired by this trick, we propose using Action Bundle that the agent predicts kk strokes at each step and the renderer renders these strokes in order. This practice encourages the exploration of the action space and action combinations. The renderer can render kk strokes simultaneously to greatly speed up the painting process.

We experimentally find that setting k=5k=5 is a good choice that significantly improves the performance and the learning speed. It’s worth noting that we modify the reward discount factor from γ\gamma to γk\gamma^{k} to keep consistency.

3.3 WGAN Reward

GAN has been widely used as a particular loss function in transfer learning, text model and image restoration , because of its great ability in measuring the distribution distance between the generated data and the target data. Wasserstein GAN (WGAN) is an improved version of the original GAN that uses the Wasserstein-l distance, also known as Earth-Mover distance. The objective of the discriminator in WGAN is defined as

where DD denotes the discriminator, ν\nu and μ\mu are the fake samples and real samples distribution. The conditional GAN training schema is used, where fake samples are pairs of a painting and its target; and real samples are two same target images as shown in Figure 5. The prerequisite of the above objective is that DD should be under the constraints of 1-Lipschitz. To achieve the constraint, we use WGAN with gradient penalty (WGAN-GP) .

We want to reduce the differences between paintings and target images as much as possible. To achieve this, we set the difference of DD scores from sts_{t} to st+1s_{t+1} using equation (1) as the reward for guiding the learning of the actor. In experiments, we find rewards derived from DD scores is better than L2L_{2} distance.

4 Network Architectures

Due to the high variability and complexity of real-world images, we use residual structures similar to ResNet-18 as the feature extractor in the actor and the critic. The actor works well with Batch Normalization (BN) , but BN can not speed up the critic learning significantly. We use WN with Translated ReLU (TReLU) on the critic to stabilize our learning. In addition, we use CoordConv as the first layer in the actor and the critic. For the discriminator, we use a network architecture similar to PatchGAN , and with WN and TReLU. The network architectures of the actor, critic and discriminator are shown in Figure 6 (a) and (b).

Following the original DDPG paper, we use the soft target network which creates a copy for the actor and critic and updating their parameters by having them slowly track the learned networks. We also apply this trick on the discriminator to improve its training stability.

Stroked-based Renderer

In this section, we introduce how to build a neural stroke renderer and use it to generate multiple types of strokes.

Using a neural network to generate strokes has two advantages. First, the neural renderer is flexible to generate any styles of strokes and is more efficient on GPU’s than most hand-crafted stroke simulators. Second, the neural renderer is differentiable and enables end-to-end training which boosts the performance of the agent.

Specifically, the neural renderer has as input a set of stroke parameters ata_{t} and outputs the rendered stroke image SS. The training samples are generated randomly using Computer Graphics rendering programs. The neural renderer can be quickly trained with supervised learning and runs on the GPU. The model-based transition dynamics st+1=trans⁡(st,at)s_{t+1}=\operatorname{trans}(s_{t},a_{t}) and the reward function r(st,at)r(s_{t},a_{t}) are differentiable. Some simple geometric trajectories like circles have simple closed-form gradients. However, in general, the discreteness of pixel position and pixel values requires continuous approximation when deriving gradients, e.g. for Bézier Curves. The approximations need to be carefully designed to not break the agent learning.

The neural renderer is a neural network consisting of several fully connected layers and convolution layers. Sub-pixel upsampling is used to increase the resolution of strokes in the network, which is a fast running operation and can eliminate the checkerboard effect. We show the network architecture of the neural renderer in Figure 6 (c).

2 Stroke Design

Strokes can be designed as a variety of curves or geometries. In general, the parameter of a stroke should include the position, shape, color, and transparency.

We design a stroke represent of quadratic Bézier curve (QBC) with thickness to simulate the effects of brushes. The shape of the Bézier curve is specified by the coordinates of control points. Formally, the stroke is defined as the following tuple:

where (x0,y0,x1,y1,x2,y2)(x_{0},y_{0},x_{1},y_{1},x_{2},y_{2}) are the coordinates of the three control points of the QBC. (r0,t0)(r0,t0), (r1,t1)(r1,t1) control the thickness and transparency of the two endpoints of the curve, respectively. (R,G,B)(R,G,B) controls the color. The formula of QBC is:

As changing stroke representation only requires changing the final stroke rendering layer, we can use neural renders with the same network structure to implement the rendering of different stroke designs.

Experiments

Four datasets are used for our experiments, including MNIST , SVHN , CelebA and ImageNet . We show that the agent has excellent performance in painting various types of real-world images.

MNIST contains 70,000 examples of hand-written digits, where 60,000 are training data, and 10,000 are testing data. Each example is a grayscale image of 28×2828\times 28 pixels.

SVHN is a real-world streetview house number image dataset, including 600,000 digits images. Each sample in the Cropped Digits set is a color image of 32×3232\times 32 pixels. We randomly sample 200,000 images for our experiments.

CelebA contains approximately 200,000 celebrity face images. The officially provided center-cropped images are used in our experiments.

ImageNet (ILSVRC2012) contains 1.2 million natural scene images, which fall into 1000 categories. The extreme diversity of ImageNet poses a grand challenge to the painting agent. We randomly sample 200,000 images that cover 1,000 categories as training data.

In our task, we aim to train an agent that can paint any images rather than only the ones in the training set. Thus, we additionally split out testing set to test the generalization ability of the trained agent. For MNIST, we use the officially defined testing set. For other datasets, we randomly split out 2,000 images as the testing set.

2 Training

We resized all images to the resolution of 128×128128\times 128 pixels before feeding the agent. With an action bundle containing 5 strokes, it takes about 2.1s2.1s to paint an image using 200 strokes on a 2.2GHz Intel Core i7 CPU. On an NVIDIA 2080Ti GPU, a 9.5×9.5\times acceleration can be achieved. The computation cost of the actor and renderer are about 554554 MFLOPs and 217217 MFLOPs respectively for painting an action bundle.

We trained the agent with 2×1052\times 10^{5} mini-batches for ImageNet and CelebA datasets, 10510^{5} mini-batches for SVHN and 2×1042\times 10^{4} mini-batches for MNIST. Adam was used for optimization, and the minibatch size was set as 96. The agent training was performed on a single GPU. It took about 40 hours for training on ImageNet and CelebA, 20 hours for SVHN and two hours for MNIST. It took 5 to 15 hours to train the neural renderer for every stroke design. The same trained renderer can be used for different agents.

At each iteration, we update the critic, actor, and discriminator in turn. All models are trained from scratch. The replay memory buffer was set to store the data of the latest 800 episodes for training the agent. Please refer to the supplemental materials for more training details.

3 Results

The images of MNIST and SVHN show simple image structures and regular contents. We train one agent that paints five strokes for images of MNIST, and another one that paints 40 strokes for images of SVHN. The example paintings are shown in Figure 3 (a) and (b). The agents can perfectly reproduce the target images.

In contrast, the images of CelebA have more complex structures and diverse contents. We train a 200-strokes agent to deal with the images of CelebA. As shown in Figure 3 (c), the paintings are quite similar to the target images although losing a certain level of details.

We train a 400-strokes agent to deal with the images of ImageNet, due to the extremely complex structures and diverse contents. As shown in Figure 3 (d), paintings are similar to the target images concerning the outline and colors of objects and backgrounds. Despite the loss of some textures, the agent still shows great power in decomposing complicated scenes into strokes and can reasonably repaint them.

We show the test loss curves of agents trained on different datasets in Figure 9.

4 Ablation Studies

In this section, we study how the components or tricks affect the performance of the agent. The control experiments are performed on CelebA.

We explore how much benefits are brought by model-based DDPG over original DDPG. Original DDPG can only essentially model the environment with observations and rewards from the environment. Besides, the high-dimensional action space also stops model-free methods from successfully dealing with the painting task. To further explore the capability of model-free methods, we improve original DDPG with a method inspired by PatchGAN. We split the images into patches before feeding the critic, then use the patch-level rewards to optimize the critic. We term this method as PatchQ. PatchQ boosts the sample efficiency and improves the performance of the agent by providing much more supervision signals in training.

4.2 Rewards

4.3 Stroke Number and Action Bundle

The stroke number for painting is critical for the final painting quality, especially for texture-rich images. We train agents that can paint 100, 200, 400 and 1000 strokes, and the testing loss curves are shown in Figure 8 (c). It’s observed that larger stroke numbers contribute to better painting quality. We show the paintings with 200-strokes and 1000-strokes in Figure 8 (e) and (f) respectively. To the best of our knowledge, few methods can handle such a large number of strokes. More strokes help reconstruct the details in the paintings.

We show test loss curves of several settings of Action Bundle in Figure 8 (b). We find that making the agent predict five strokes in one bundle achieves the best performance. We conjecture that increasing strokes number in one bundle helps the agent to build long-term plans as there will be fewer rounds of decision, even though it will increase the difficulty in a single round of decision. Thus, to achieve a trade-off, a few strokes in one bundle is a good setting for the agent. Experiments determine that setting five actions in an Action Bundle is optimal in our setting for the painting task.

4.4 Alternative Stroke Representations

Besides the QBC, we find alternative stroke representations can also be well mastered by the agent, including straight strokes, circles, and triangles. We train one neural renderer for each stroke representation. The paintings with these renderers are shown in Figure 10. The QBC strokes produce excellent visual effects. Meanwhile, other stroke designs create different artistic effects. Although with different styles, the paintings still resemble the target images. This shows that our network architecture generalizes to other choices of stroke designs.

Also, by restricting the transparency of strokes, we can get paintings with different stroke effects, such as ink painting and oil painting as shown in Figure 7 (c).

Conclusion

In this paper, we train agents that decompose the target image into an ordered sequence of strokes in a fashion mimicking human painting processes on canvases. The training is based on the Deep Reinforcement Learning framework, which encourages the agent to make long-term plans for sequential stroke-based painting. In addition, we build a differentiable neural renderer to render the strokes, which allows using model-based DRL algorithms to further improve the quality of recreated images. The learned agent can predict hundreds or even thousands of strokes to generate a vivid painting. Experimental results show that our model can handle multiple types of target images and achieve good performance in painting real-world images like human portraits and texture-rich natural scenes.

References

Appendix

The network structure diagrams are shown in Figure 11, 12, 13 and 14 , where FC refers to a fully-connected layer, Conv is a convolution layer. The hyperparameters used in training are listed as much as possible in Table 1 and 2. All ReLU activations between the layers have been omitted for brevity.