CLIP-NeRF: Text-and-Image Driven Manipulation of Neural Radiance Fields

Can Wang, Menglei Chai, Mingming He, Dongdong Chen, Jing Liao

Introduction

With the explosive growth of 3D assets, the demand for manipulating 3D content to achieve versatile re-creation is rising rapidly. While most existing 3D editing methods operate on explicit 3D representations zwicker2002pointshop; ju2007editing; fried2019text, the recent advances of implicit volumetric representations in capturing and rendering dedicated 3D structures park2019deepsdf; mildenhall2020nerf; riegler2020free; jiang2020local; genova2020local; kar2017learning have motivated the research to benefit the manipulation from such representations. Among these works, neural radiance fields (NeRF) mildenhall2020nerf utilize a volume rendering technique to render neural implicit representations for high-quality novel view synthesis, providing an ideal representation for 3D content.

Editing NeRF (e.g., deforming the shape or changing the appearance color), however, is an extremely challenging task. First, since NeRF is an implicit function optimized per scene, we cannot directly edit the shape using the intuitively tools for the explicit representations schmidt2016state; wang20193dn; wang2020pixel2mesh; uy2020deformation. Second, unlike image manipulation where the single-view information is enough to guide the editing li2020manigan; xia2021tedigan; xu2021text, the multi-view dependency of NeRF makes the manipulation way more difficult to control without the multi-view information. More recent works propose conditional NeRF schwarz2020graf, which trains NeRF on one category of shapes and enables manipulation via latent space interpolations utilizing the pre-trained models. Based on the conditional NeRF, EditNeRF liu2021editing takes the first step to edit the shape and color of NeRF given user scribbles. However, due to its limited capacity in shape manipulation, only adding or removing local parts of the object is allows. In addition to achieving more compelling and complicated manipulation, we seek to edit NeRF in more intuitive ways, such as using a text prompt or a single reference image.

In this paper, we explore how to individually manipulate the shape and the appearance of NeRF based on a text prompt or a reference image in a unified framework. Our framework is built on a novel disentangled conditional NeRF architecture, which is controlled by the latent space disentangled into a shape code and an appearance code. The shape code guides the learning of a deformation field to warp the volume to a new geometry, while the appearance code allows controlling the emitted color of volumetric rendering. Based on our disentangled NeRF model, we take advantage of the recently proposed Contrastive Language-Image Pre-training (CLIP) model radford2021learning to learn two code mappers, which map CLIP features to the latent space to manipulate the shape or appearance code. Specifically, given a text prompt or an exemplar image as our condition, we extract the features using the pre-trained CLIP model, feed the features into the code mappers, and yield local displacements in the latent space to edit the shape and appearance codes to reflect the edit. We design the CLIP-based loss to enforce the CLIP space consistency between the input constraint and the output renderings, thus supporting high-resolution NeRF manipulation. Additionally, we propose an optimization-based method for editing a real image by inversely optimizing its shape and appearance codes.

To sum up, we make the following contributions:

We present the first text-and-image-driven manipulation method for NeRF, using a unified framework to provide users with flexible control over 3D content using either a text prompt or an exemplar image.

We design a disentangled conditional NeRF architecture by introducing a shape code to deform the volumetric field and an appearance code to control the emitted colors.

Our feedforward code mappers enable the fast inference for editing different objects in the same category compared to the optimization-based editing method liu2021editing.

We propose an inversion method to infer the shape and appearance codes from a real image, allowing editing the shape and appearance of the existing data.

Related Work

NeRF and NeRF Editing. The past few years have witnessed tremendous progress in the implicit representation of 3D models with neural networks park2019deepsdf; mildenhall2020nerf; riegler2020free; jiang2020local; genova2020local; kar2017learning. Among them, NeRF mildenhall2020nerf is a representative one, which encodes a continuous volume representation of shape and view-dependent appearance in the weights of an MLP network. NeRF has been gaining more and more popularity because of its strong capability in capturing high-resolution geometry and rendering photo-realistically novel views. The success of NeRF has also inspired many follow-up works that extend the NeRF to dynamic scenes park2020deformable; pumarola2021d; gafni2021dynamic; tretschk2021non, relighting boss2021nerd; srinivasan2021nerv, generative models schwarz2020graf; niemeyer2021giraffe; chan2021pi; jang2021codenerf, etc. Furthermore, DietNeRF jain2021putting designs a CLIP semantic consistency loss to improve few-shot NeRF and presents impressive results, and GRAF schwarz2020graf first adopts shape and appearance codes to conditionally synthesize NeRF, which inspires our adversarial training.

Despite the above success, a 3D model with NeRF representation is very unintuitive and difficult to edit since it is represented by millions of network parameters. To address this problem, the pioneering work EditNeRF liu2021editing defines a conditional NeRF, where the 3D object encoded by NeRF is conditioned on a shape code and an appearance code. By optimizing the adjustment to these two latent codes, user edits on shape and appearance color can be achieved. However, this method has limited capacity in shape manipulation as it only supports adding or removing local parts of the object. Also, the editing process of EditNeRF liu2021editing is slow because of its iterative optimization nature. Compared to EditNeRF liu2021editing, our method is different in three aspects. First, our method gives more freedom in shape manipulation and supports global deformation. Second, by learning two feed-forward networks mapping user edits to the latent codes, our method allows fast inference for the interactive editing. Moreover, different from the user scribbles used in EditNeRF liu2021editing, we introduce two intuitive ways to NeRF editing: using either a short text prompt or an exemplar image, which are more friendly to novice users.

CLIP-Driven Image Generation and Manipulation. An important building block of our work is CLIP radford2021learning which connects texts and images by bringing them closer in a shared latent space, under a contrastive learning manner. Powered by the CLIP model, some text-driven image generation and manipulation methods are proposed. Perez perez2021imagesfromprompts combines CLIP and StyleGAN karras2020analyzing; karras2019style to synthesize images by optimizing the latent code of a pre-trained StyleGAN according to a textual condition defined in the CLIP space. Instead of generating images from scratch, StyleCLIP patashnik2021styleclip introduces a text-based interface for StyleGAN to allow manipulations of real images with text prompts. Besides applying CLIP to GAN models, DiffusionCLIP kim2021diffusionclip combines a diffusion model song2019generative with CLIP to conduct a text-driven image manipulation. It achieves a comparable performance to that of GAN-based image manipulation methods, with the advantage of great mode coverage and training stability. However, all these methods only explore the text-guidance ability of CLIP, whereas our method unifies both text-and-image driven manipulations in a single model by fully exploiting the power of CLIP. Further, these methods are limited to image manipulation and fail to encourage multi-view consistency due to the lack of 3D information. In contrast, our model combines NeRF with CLIP, thus allowing editing 3D models in a view consistent way.

Method

In this section, we start with the general formulation of conditional NeRF (§ 3.1) as a 3D generative model conditioned by shape and appearance codes. We then present our disentangled conditional NeRF model (§ 3.2), which is able to individually control the shape and appearance manipulation. Next, we introduce our framework on leveraging the multi-modal power of CLIP for driving NeRF manipulation (§ 3.3) using both text prompts or image exemplars, and the training strategy (§ 3.4). Finally, we propose an inversion method (§ 3.5) to allow editing a real image by a novel latent optimization approach on shape and appearance codes.

Built upon the original per-scene NeRF, conditional NeRF servers as a generative model for a particular object category, conditioned on the latent vectors that dedicatedly control shape and appearance. Specifically, conditional NeRF is represented as a continuous volumetric function Fθ\mathcal{F}_{\theta} that maps a 5D coordinate (a spatial position x(x,y,z){\bm{x}}(x,y,z) and a view direction v(ϕ,θ){\bm{v}}(\phi,\theta)), together with a shape code zs{\bm{z}}_{s} and an appearance code za{\bm{z}}_{a}, to a volumetric density σ\sigma and a view-dependent radiance c(r,g,b){\bm{c}}(r,g,b), parametrized by a multi-layer perceptron (MLP). A trivial formulation Fθ′(⋅)\mathcal{F}^{\prime}_{\theta}(\cdot) of conditional NeRF can be:

where ⊕\oplus is the concatenation operator.

where k∈{0,…,2m−1}k\in\{0,\ldots,2m-1\} and mm is a hyper-parameter that controls the total number of frequency bands.

2 Disentangled Conditional NeRF

The aforementioned conditional NeRF does introduce conditional generation capability to the NeRF architecture. However, this trivial formulation Fθ′\mathcal{F}^{\prime}_{\theta} (Eq. 1) suffers from mutual intervention between shape and appearance conditions, e.g., manipulating the shape code could also cause color changes. In observation of this issue, we propose our disentangled conditional NeRF architecture to achieve individual control over both shape and appearance by properly disentangling the conditioning mechanism.

Conditional Shape Deformation. Rather than directly concatenating the latent shape code to the encoded position feature, we propose to formulate the shape conditioning through explicit volumetric deformation to the input position. This conditional shape deformation not only improves the robustness of the manipulation and preserves the original shape details as much as possible by regularizing the output shape to be a smooth deformation of the base shape, but more importantly also completely isolates the shape condition from affecting the appearance.

Deferred Appearance Conditioning. In NeRF, the density is predicted first as a function of position and the radiance is then predicted from both position and view direction. Similar to Graf schwarz2020graf and EditNeRF liu2021editing, we also defer the appearance conditioning to concatenate the appearance code with the view direction as the input to the radiance prediction network, which allows manipulating the appearance without touching the shape information, i.e., density.

Overall, as illustrated in Fig. 1, our disentangled conditional NeRF Fθ(⋅)\mathcal{F}_{\theta}(\cdot) is defined as:

And for the simplicity of notation, we use Fθ(v,zs,za)={Fθ(x,v,zs,za)∣x∈R}\mathcal{F}_{\theta}({\bm{v}},{\bm{z}}_{s},{\bm{z}}_{a})=\big\{\mathcal{F}_{\theta}({\bm{x}},{\bm{v}},{\bm{z}}_{s},{\bm{z}}_{a})\mid{\bm{x}}\in\mathbf{R}\big\} to denote the rendering of the whole image with viewport R\mathbf{R}.

3 CLIP-Driven Manipulation

With our disentangled conditional NeRF (Eq. 4) as a generator, we now introduce how we integrate the CLIP model into the pipeline to achieve text-driven manipulation on both shape and appearance.

To avoid optimizing both shape and appearance codes for each target sample, which tends to be versatile and time-consuming, we take a feed-forward approach to directly update the condition codes from the input text prompt. Specifically, given an input text prompt of t{\bm{t}} and the initial shape/appearance code of zs′{\bm{z}}_{s}^{\prime}/za′{\bm{z}}_{a}^{\prime}, we train a shape mapper Ms\mathcal{M}_{s} and an appearance mapper Ma\mathcal{M}_{a} to update the codes as:

where E^t(⋅)\hat{\mathcal{E}}_{t}(\cdot) is the pre-trained CLIP text encoder that projects the text to the CLIP embedded feature space and both mappers map this CLIP embedding to displacement vectors that update the original shape and appearance codes.

In addition, given that CLIP includes an image encoder and a text encoder mapping to a joint embedding space, we define a cross-modal CLIP distance function DCLIP(⋅,⋅)D_{\text{CLIP}}(\cdot,\cdot) to measure the embedding similarity between the input text and a rendered image patch:

where E^i(⋅)\hat{\mathcal{E}}_{i}(\cdot) and E^t(⋅)\hat{\mathcal{E}}_{t}(\cdot) are the pre-trained CLIP image and text encoders, I\mathbf{I} and t{\bm{t}} are the input image patch and text, and ⟨⋅,⋅⟩\langle\cdot,\cdot\rangle is the cosine similarity operator.

Without loss of generality, here we assume that the manipulation control comes from a text prompt t{\bm{t}}. However, our distance can also be extended to measure similarity between two images or two text prompts. Thus, our framework naturally supports editing with an image exemplar by trivially replacing the text prompt with this exemplar in aforementioned equations.

Discussion. To perform NeRF manipulation with image-level CLIP model, a natural question is whether the CLIP feature is stable across different viewpoints and whether it can distinguish object differences. To evaluate this, we randomly select two objects (e.g., an SUV and a jeep) and measure the pairwise CLIP-space cosine distances between 1) different views of a same object, and 2) different objects in a same view. As shown in Fig. 2, we find the distance is more sensitive to small object difference than large view variations. This suggests that a pre-trained CLIP model has the ability to support view-consistency representations for 3D-aware applications. A similar observation is found by DietNeRF jain2021putting and applied in 3D reconstruction.

4 Training Strategy

Our pipeline is trained in two stages: we first train the disentangled conditional NeRF including the conditional NeRF generator and the deformation network; then we fix the weights of the generator and train the CLIP manipulation parts including both the shape and appearance mappers.

Disentangled Conditional NeRF. Our conditional NeRF generator Fθ\mathcal{F}_{\theta} is trained together with the deformation network using a non-saturating GAN objective mescheder2018training with the discriminator D\mathcal{D}, where f(x)=−log⁡(1+exp⁡(−x))f(x)=-\log\big(1+\exp(-x)\big) and λr\lambda_{r} is the regularization weight. Assuming that real images I\mathbf{I} form the training data distribution of dd, we randomly sample the shape code zs{\bm{z}}_{s}, the appearance code za{\bm{z}}_{a}, and the camera pose from Zs\mathcal{Z}_{s}, Za\mathcal{Z}_{a}, and Zv\mathcal{Z}_{v}, respectively, where Zs\mathcal{Z}_{s} and Za\mathcal{Z}_{a} are the normal distribution, and Zv\mathcal{Z}_{v} is the upper hemisphere of the camera coordinate system.

CLIP Manipulation Mappers. We use pre-trained NeRF generator Fθ\mathcal{F}_{\theta}, CLIP encoders {E^t,E^i}\{\hat{\mathcal{E}}_{t},\hat{\mathcal{E}}_{i}\}, and the discriminator D\mathcal{D} to train the CLIP shape mapper Ms\mathcal{M}_{s} and appearance mapper Ma\mathcal{M}_{a}. All network weights, except the mappers, are fixed, denoted as {⋅^}\{\hat{\cdot}\}. Similar to the first stage, we randomly sample the shape code zs{\bm{z}}_{s}, the appearance code za{\bm{z}}_{a}, and the camera pose v{\bm{v}} from their respective distributions. In addition, we sample the text prompt t{\bm{t}} from a pre-defined text library T\mathbf{T}. By using our CLIP distance DCLIPD_{\text{CLIP}} (Eq. 6) with weight λc\lambda_{c}, we train the mappers with the following losses:

5 Inverse Manipulation

The manipulation pipeline we have introduced so far works on an initial sample with known conditions including the shape and appearance codes. To apply the manipulation to an input image Ir\mathbf{I}_{r} belonging to the same training category, the key is to first optimize all generation conditions to inversely project the image to the generation manifold, similar to the latent image manipulation methods abdal2019image2stylegan; abdal2020image2stylegan++; gu2020image; pan2021exploiting. Following the EM algorithm dempster1977maximum, we design an iterative method to alternatively optimize the shape code zs{\bm{z}}_{s}, the appearance code za{\bm{z}}_{a}, and the camera v{\bm{v}}.

To be specific, during each iteration, we first optimize v{\bm{v}} while keeping zs{\bm{z}}_{s} and za{\bm{z}}_{a} fixed using the following loss:

We then update the shape code by minimizing:

where za{\bm{z}}_{a} and v{\bm{v}} are fixed, zn{\bm{z}}_{n} is a random standard Gaussian noise vector sampled in each step to improve the optimization robustness, and λn\lambda_{n} linearly decays from 11 to 00 through the whole optimization iterations.

The appearance code is updated in a similar manner:

Experiments

Datasets. We evaluate our method on two public datasets: Photoshapes park2018photoshape; schwarz2020graf with 150K chairs rendered at 128×\times128 following the rendering protocol of oechsle2020learning and Carla with 10K cars rendered at 256×\times256 using the Driving simulator dosovitskiy2017carla; schwarz2020graf. Each object is rendered in a random view without providing any camera pose parameters.

We compare with pioneering work in NeRF editing, EditNeRF liu2021editing on the editing of shape and appearance color of both datasets in Fig. 3. For the Photoshapes dataset, EditNeRF is trained using 600 instances with 40 views per instance while ours uses only one view. For the Carla dataset, EditNeRF uses 10K cars with a single view per instance, same as ours. Besides, camera pose parameters are required during training of EditNeRF but unknown for us.

We first compare the capability and performance between EditNeRF and our method. For the color editing (Fig. 3-(a)), EditNeRF requires the user to select a target color and draw coarse scribbles on a local region. With the foreground and background masks created by the coarse scribbles, EditNeRF performs the appearance editing by optimizing the appearance code and a conditional NeRF to achieve the target color. We observe that unnatural color effects appear on the edited results of EditNeRF (e.g. discontinuity on car doors), and the generated color is not completely faithful to the target color. In contrast, we allow the user to change the color more simply by providing a text prompt and our method produces more natural editing results (Fig. 3-(b)). For the shaping editing (Fig. 3-(c)), EditNeRF can only support local shape editing, such as shape part removal. Given the user’s editing scribbles which for example indicate to remove a leg of a chair (in the red rectangles), EditNeRF optimizes a few layers in the network to fit the shape in the input view but it cannot ensure successful propagation to unseen views (in the blue rectangle) and keep the structure of other parts intact (in the green rectangle). Compared to it, our method supports a large degree of shape deformation and generalizes well to unseen views (Fig. 3-(d)). Besides, EditNeRF, as an optimized-based method, takes a large amount of time for the optimization, while our feedforward code mappers achieve much faster inference of the target shape and appearance (Table 1).

To quantitatively evaluate how good the image quality is preserved after editing, we calculate the FID scores of 2K testing images before and after editing. Due to training with 40 views per instance, EditNeRF shows better reconstruction before editing on the chair dataset, but its editing notably degrades the image quality while our method ensures comparable quality before and after manipulation. On the car dataset, the performance of EditNeRF significantly drops because of only one view per instance used in training. Trained under the same setting, our model improves reconstruction quality by a large margin and well preserves the quality during editing. Since EditNeRF requires user scribbles for shape editing and is difficult to generate a large set of results with random conditions, we exclude it in the comparison on shape editing while our method performs equally well regardless of editing shape or color.

We also compare with EditNeRF in the inversion results (Fig. 5). EditNeRF infers the shape and appearance codes by fine-tuning the condition NeRF using the standard NeRF photometric loss. Our optimized-based inversion method outperforms by benefiting from CLIP’s ability to provide multi-view consistency representations (more discussions in Section 3.3 and ablation in 4.2).

2 Ablation Study

We evaluate our model w/ and w/o the disentangled design (Section 3.2). In Fig. 4, the model trained without the conditional shape deformation network (i.e. w/o disen.) frequently introduces color changes when performing shape editing. In contrast, our disentangled conditional NeRF achieves individual shape control because the conditional shape deformation network is able to isolate the shape condition from the appearance control and deform the base volume field to generate new objects without affecting the appearance. Also, since the deformation network implicitly enforces regularization of the generated shape, the resulting quality is further improved as shown in Table 4.2.

We conduct another ablation study in Fig. 5 to evaluate our inversion optimization method. The baseline method (w/o CLIP) only computes the standard NeRF photometric loss between the output and a single image. Its result quality is limited due to the difficulty in inferring a complete 3D NeRF model from a single view. As discussed in Section 3.3, CLIP has the capability to produce robust pose-invariant features. Therefore, our inversion method introduces a CLIP constraint during optimization and achieves better inversion results thanks to the CLIP prior.

3 CLIP-Driven Manipulation

Our method supports editing of object shape or appearance using text. When manipulating the shape, we keep the appearance code unchanged and the same applies to appearance editing. Fig. 6 demonstrates diverse editing results. Note that when the car with a light color is deformed to a sports car, its color may become darker. But it is not a failure case, as the colors of all sports cars in the Carla dataset are inherently intenser. Besides, we find that our method naturally preserves the shading when changing the appearance color. When editing the chair shape, if the user’s input text is highly relevant to the source shape−-for example, the source chair is a wood chair, and the user also wants a ’wood chair’−-the result will be slightly different from the source. During the color editing, our method guarantees that the shape is completely preserved. Our method also supports exemplar-based manipulation by providing a real target image instead of a text prompt. We present various exemplar-guided shape and appearance editing results in Fig. 7. Our method achieves semantic-precise and individual control of shape and appearance referring to the exemplar image.

4 Real Image Manipulation

To evaluate the generalization ability of our model in processing a single real image that does not exist in our training set, we experiment with the real image by inverting it to a shape code and an appearance code and then applying them to edit. We show the inverted and edited results in Fig. 8. We observe that inverting the chair is much more challenging than inverting a car due to the chair’s delicate structures, such as the wheels of the office chair. However, even the office chair is not perfectly reconstructed, the editing ability of our method is not affected. Our method still ensures accurate editing in shape and appearance.

5 User Study

We conduct a user study to evaluate the perceptual quality and accuracy of the editing results. We include 20 questions in the study, each question with 5 results of cars or chairs generated by 5 randomly selected text prompts or 5 randomly selected exemplars. We randomly shuffle the results and give users unlimited time to match each result with the correct text or image. We collect answers from 23 participates and report the matching accuracy rate in Table 3. Our method, in more than 80% cases, succeeds in editing objects exactly corresponding to the description given by the text or the exemplar.

Extended Discussions

Continuous Manipulation. Our method supports editing both shape and appearance, given a single text prompt or an examplar. This can be achieved by continuously editing the shape and appearance, i.e., first editing the shape and then editing the color and vice versa. We show results in Fig. 9. This provides a user-friendly way for editing when users want to edit both the shape and appearance indicated by a single text description or an exemplar.

Fine-grained appearance manipulation within a same color category. Though our method cannot handle fine-grained local parts shape and appearance edits as stated in the limitation, it supports fine-grained appearance manipulation at a whole object level, as shown in Fig. 10. Our method enjoys achieving various editing results within a same color category. Without loss of generality, we show various editing results related to the color blue.

Scaling along Editing Direction. From equation 13, our code mappers provide manipulation directions Δzs=(E^t(t))\Delta z_{s}=\big(\hat{\mathcal{E}}_{t}({\bm{t}})\big) and Δza=Ma(E^t(t))\Delta z_{a}=\mathcal{M}_{a}\big(\hat{\mathcal{E}}_{t}({\bm{t}})\big) in the latent space for shape and appearance editing. We can scale along the editing direction to obtain gradually editing results through the following equation:

where ss is the scalar. This scaled scheme also supports directions learned from examplars. We show scaled manipulation results in Fig. 12. The manipulation effect becomes stronger as the scalar ss increases.

Interpolation. As shown in Fig. 13, the shape and appearance latent space supports interpolation between two latent codes z1=(zs1,za1){\bm{z}}^{1}=({\bm{z}}_{s}^{1},{\bm{z}}_{a}^{1}) and z2=(zs2,za2){\bm{z}}^{2}=({\bm{z}}_{s}^{2},{\bm{z}}_{a}^{2}). Given an interpolation ratio r{\bm{r}}, we define the interpolated latent code zinter{\bm{z}}_{inter} as zinter=z2×r+z1×(1−r){\bm{z}}_{inter}={\bm{z}}^{2}\times{\bm{r}}+{\bm{z}}^{1}\times(1-{\bm{r}}), while r{\bm{r}} ranges from 0 to 1.0 with a step 0.1. Then we can obtain the interpolated result using zinter{\bm{z}}_{inter}.

Necessity of latent space. Our method performs shape and appearance edits on the latent space of a conditional NeRF model with our designed CLIP constraints. A question arises whether the latent space is necessary, i.e., is it possible to edit the shape and appearance of a single NeRF model mildenhall2020nerf directly rather than a conditional NeRF? We first evaluate our designed CLIP loss on appearance editing in Fig. 14. Given a pre-trained NeRF model on the LLFF dataset mildenhall2019local, we fix the density-related layers and finetune the color-related layers of NeRF with our CLIP loss. We also use the patch-based ray samplar while calculating our CLIP loss. Our CLIP constraint succeeds in editing the color of a single NeRF model without any ground truth. However, we fail to achieve satisfying results while editing the shape of a single NeRF with our deformation network conditioned by a text prompt. This may be because the CLIP loss is still not strong and compact enough to deform the shape without the latent space constraint. We think it is an interesting problem to explore in the future.

Supplementary Video

We provide a supplementary video with a real-time demo and more visual results rendered in multiple views. We highly recommend watching our supplementary video to observe the user-friendliness and view-consistency that our method can achieve in both shape and color editing.

Conclusion

We present the first text-and-image driven manipulation method for NeRF by designing a unified framework to provide users with flexible control over 3D content using either a text prompt or an exemplar image. We design a disentangled conditional NeRF architecture that allows disentangling shape and appearance while editing an object, and two feedforward code mappers enable fast inference for editing different objects. Further, we proposed an inversion method to infer the shape and appearance codes from a real image, allowing editing the existing data.

Limitations. We evaluate our approach by extensive experiments on various text prompts and exemplar images and provide an intuitive editing interface for interactive editing. However, our method cannot handle fine-grained and out-of-domain shape and appearance edits as shown in Fig. 11, due to the limited expressive ability of the latent space and the pre-trained CLIP. This may be alleviated by adding more various training data.

References