Re-imagine the Negative Prompt Algorithm: Transform 2D Diffusion into 3D, alleviate Janus problem and Beyond

Mohammadreza Armandpour, Ali Sadeghian, Huangjie Zheng, Amir Sadeghian, Mingyuan Zhou

Introduction

Advancements in generating images using diffusion models from text have shown remarkable capabilities in producing a wide range of creative images from unstructured text inputs . However, research has found that the generated images may not always accurately represent the intended meaning of the original text prompt .

Generating satisfactory images that semantically match the text query is challenging, as it requires textual concepts to match the images at a grounded level. However, due to the difficulty of obtaining such a fine-grained annotation, current text-to-image models have difficulty fully understanding the relationship between text and images. Therefore, they are inclined to generate images like high-frequent text-image pairs in the datasets, where we can observe that the generated images are missing requested or containing undesired attributes . Most of the recent works focus on adding back the missing objects or attributes to existing content to edit images based on a well-designed main text prompt . However, limited of them study how to remove redundant attributes, or force the model NOT to have an unwanted object using negative prompts , which is the main goal of our paper.

We start this paper by showing the shortcoming of the current negative prompt algorithm. After our initial investigation, we realized the current implementation of using negative prompts could produce unsatisfactory results when there is an overlap between the main prompt and the negative ones, as shown in the examples in Figure 2. To address the above problem, we propose Perp-Neg algorithm, which does not require any training and can readily be applied to a pre-trained diffusion model. We refer to our method as Perp-Neg since it employs the perpendicular score estimated by the denoiser for the negative prompt. More specifically, Perp-Neg limits the direction of denoising, guided by the negative prompt to be always perpendicular to the direction of the main prompt. In this way, the model is able to eliminate the undesired perspectives in the negative prompts without changing the main semantics, as illustrated in Figure 2.

Furthermore, We extend Perp-Neg to DreamFusion, a state-of-the-art text-to-3D model, and show how Perp-Neg can alleviate its Janus problem, which refers to the case that a 3D-generated object inaccurately shows the canonical view of the object from several viewpoints, as shown in the left column of Figure 1. Recent studies have considered that the main cause of the Janus problem is the failure of the pre-trained 2D diffusion model in following the view instruction provided in the prompt . Therefore, we first, in 2D, show quantitatively and qualitatively how our algorithm can significantly improve the view fidelity of a pretrained diffusion model. We also explore how Perp-Neg can be employed for effective interpolation between two views of an object in 2D as it is needed for 3D cases, as illustrated in Figure 3. Then we integrate Perp-Neg in Stable DreamFusion and show how it can alleviate the Janus problem.

Our contributions can be summarized as follows:

We find the limitations of the current negative prompt implementation which is susceptible to the overlap between a positive and a negative prompt.

We propose Perp-Neg, a sampling algorithm for text-to-image diffusion models to eliminate undesired attributes indicated by the negative prompt while preserving the main concept, without any training needed.

Our experiments quantitatively and qualitatively demonstrate that Perp-Neg significantly improves diffusion model prompt fidelity in view generation.

By enhancing the 2D diffusion model in following the view instruction, we mitigate the Janus problem in text-to-3D generation tasks.

Perp-Neg: Novel negative prompt algorithm

Diffusion Models: Diffusion-based (also known as score-matching) models is a family of generative models that employ a forward process and a reverse process to iteratively corrupt and generate the data within TT steps. Specifically, denoting q(x0)q({\bm{x}}_{0}) as the data distribution and p(xT)p({\bm{x}}_{T}) as the generative prior, such two processes can be modeled as the following:

One of the most appealing attributes of diffusion models is that any intermediate step of the forward process and every single step in the reverse process can be modeled as a Gaussian distribution like formulated in :

where {αt}t=1T\{\alpha_{t}\}_{t=1}^{T} and {σt}t=1T\{\sigma_{t}\}_{t=1}^{T} can be explicitly calculated with a pre-defined variance schedule {βt}t=1T\{\beta_{t}\}_{t=1}^{T}. Moreover, the generator μθ(⋅)\mu_{\bm{\theta}}(\cdot) is a linear combination of xt{\bm{x}}_{t} and a trainable generator ϵθ{\bm{\epsilon}}_{\bm{\theta}} that predicts the noise in xt{\bm{x}}_{t}, which is usually optimized with a simple weighted noise prediction loss

with w(t)w(t) as the weight that depends on the timestep tt that is uniformly drawn from {1,...,T}\{1,...,T\}.

Text-to-Image Diffusion Models and Composing Diffusion Model: Recent works have shown the success of leveraging the power of diffusion models, where large-scale models are able to be trained on extremely large text-image paired datasets by modeling with the loss function in Equation 3 (or its variants) , with the text prompt c{\bm{c}} often encoded with a pre-trained large language model . To generate photo-realistic images given text prompts, the diffusion models can further take advantage of classifier guidance or classifier-free guidance to improve the image quality. Especially, in the context of text-to-image generation, classifier-free guidance is more widely used, which is usually expressed as a linear interpolation between the conditional and unconditional prediction ϵ^θ(xt,t,c)=(1+τ)ϵθ(xt,t,c)−τϵθ(xt,t)\hat{\bm{\epsilon}}_{\bm{\theta}}({\bm{x}}_{t},t,{\bm{c}})=(1+\tau){\bm{\epsilon}}_{\bm{\theta}}({\bm{x}}_{t},t,{\bm{c}})-\tau{\bm{\epsilon}}_{\bm{\theta}}({\bm{x}}_{t},t) at each timestep tt with a guidance scale parameter τ\tau.

When the prompt becomes complex, the model may fail to understand some key elements in the query prompt and create undesired images. To handle complex textual information, proposes composing diffusion models to factorize the text prompts into a set of text prompts, i.e., c ⁣ ⁣= ⁣ ⁣{c1,...cn}{\bm{c}}\!\!=\!\!\{{\bm{c}}_{1},...{\bm{c}}_{n}\}, and model the conditional distribution as

By applying Bayes rule, we have p(ci∣x)∝p(x∣ci)p(x)p({\bm{c}}_{i}|{\bm{x}})\propto\frac{p({\bm{x}}|{\bm{c}}_{i})}{p({\bm{x}})} and

Note that pθ(x∣ci)p_{{\bm{\theta}}}({\bm{x}}|{\bm{c}}_{i}) and pθ(x)p_{{\bm{\theta}}}({\bm{x}}) respectively correspond to ϵθ(xt,t,ci){\bm{\epsilon}}_{\bm{\theta}}({\bm{x}}_{t},t,{\bm{c}}_{i}) and ϵθ(xt,t){\bm{\epsilon}}_{\bm{\theta}}({\bm{x}}_{t},t) modeled by the diffusion model. Putting them together yields a composed noise predictor, as shown in :

where wneg>0w_{\text{neg}}>0 is a weight function depending on τ\tau and β\beta, denoting the scale for the concept negation.

2 Perpendicular gradient sampling

Although proposes to decompose the text condition into a set of positive and negative prompts in order to help the model handle complex textual inputs, the proposed method assumes these conditional prompts are independent of each other, which requires careful design of the prompts or maybe too ideal to realize in practice. For simplicity of presentation, below we present the overlap problem with the case of fusing two prompts, i.e., the main prompt c1{\bm{c}}_{1} and an additional prompt c2{\bm{c}}_{2}. Without loss of generality, this problem can also be generalized to the case where the main prompt is combined with a series of prompts as {c1,...,cn}\{{\bm{c}}_{1},...,{\bm{c}}_{n}\}. To illustrate the problem, we first re-write the relation in Equation 4:

When c1{\bm{c}}_{1} and c2{\bm{c}}_{2} are conditional independent given x{\bm{x}}, the ratio R(c1,c2)=pθ(c1,c2∣x)pθ(c1∣x)pθ(c2∣x) ⁣= ⁣1\mathcal{R}({\bm{c}}_{1},{\bm{c}}_{2})=\frac{p_{{\bm{\theta}}}({\bm{c}}_{1},{\bm{c}}_{2}|{\bm{x}})}{p_{{\bm{\theta}}}({\bm{c}}_{1}|{\bm{x}})p_{{\bm{\theta}}}({\bm{c}}_{2}|{\bm{x}})}\!=\!1 and this term can be ignored. However, in practice, the input text prompts can barely be independent when we need to specify the desired attributes of the image, such as style, content, and their relations. When c1{\bm{c}}_{1} and c2{\bm{c}}_{2} have an overlap in their semantics, simply fusing the concepts could be harmful and result in undesired results, especially in the case of concept negation, as shown in Figure 2. In the second row of images, we can clearly observe the key concepts requested in the main text prompt (respectively “armchair”, “sunglasses”, “crown”, and “horse”) are removed when those concepts appear in the negative prompts. This important observation motivates us to rethink the concept composing process and propose the use of a perpendicular gradient in the sampling, which is described in the following section.

2.2 Perpendicular gradient

Recall when c1{\bm{c}}_{1} and c2{\bm{c}}_{2} are independent, both of them possess a denoising score component

and we can directly fuse these denoising scores as done in Equation 6. However, from the above section, when c1{\bm{c}}_{1} and c2{\bm{c}}_{2} overlap, we cannot directly fuse the denoising components together, which motivates us to seek the independent component of c2{\bm{c}}_{2} to ensure the fused denoising score does not hurt the semantics in c1{\bm{c}}_{1}.

Considering the geometrical interpretation of ϵθi{\bm{\epsilon}}_{\bm{\theta}}^{i} indicates the gradient that the generative model should denoise to produce the final images, a natural solution is to find the perpendicular gradient of ϵθ1{\bm{\epsilon}}_{\bm{\theta}}^{1} as the independent component of ϵθ2{\bm{\epsilon}}_{\bm{\theta}}^{2}. Therefore, we now re-formulate Equation 6 and define the Perp-Neg sampler for c1{\bm{c}}_{1} and c2{\bm{c}}_{2} as

where ⟨,⟩\langle,\rangle denotes the vectorial inner product, w1w_{1} and w2w_{2} define the weights for each component, and ⟨ϵθ1,ϵθ2⟩∥ϵθ1∥2\frac{\langle{\bm{\epsilon}}_{\bm{\theta}}^{1},{\bm{\epsilon}}_{\bm{\theta}}^{2}\rangle}{\|{\bm{\epsilon}}_{\bm{\theta}}^{1}\|^{2}} defines the projection function to find the most correlated component of c2{\bm{c}}_{2} to c1{\bm{c}}_{1}.

Note that although the proposed perpendicular gradient sampler is applicable for both positive text prompts and negative prompts, we find in the case of concept conjunction, the positive prompts can be designed to be independent of the main prompt in an easier way, as we are creating new details in complementary to the main concept. However, in the case of concept negation, it is more frequent to observe the negative prompts have overlap with the main text prompt. Compared to the sampler in Equation 6, the most important property of the perpendicular gradient is that the component of ϵθ1{\bm{\epsilon}}_{\bm{\theta}}^{1} won’t be affected by the additional prompt. Imagine the case where ϵθ1=ϵθ2{\bm{\epsilon}}_{\bm{\theta}}^{1}={\bm{\epsilon}}_{\bm{\theta}}^{2}, using Equation 6, the denoising gradient becomes zero if we also set w1=−w2w_{1}=-w_{2}, which might fail the generation. However, using perpendicular gradient in Equation 7 could still preserve the main component ϵθ1{\bm{\epsilon}}_{\bm{\theta}}^{1}. Below we mainly discuss the case of using perpendicular gradient sampling to handle the negative prompts and introduce Perp-Neg algorithm.

2.3 Perp-Neg algorithm

2D diffusion model for 3D generation

Background: Since 2D diffusion models not only provide samples of density but also allow calculating the derivate of data density likelihood. There are several seminal works that use the latter advantage to uplift a pretrained 2D diffusion and make it a 3D generative model. The main idea behind all these methods is to optimize a 3D scene representation of an object (e.g.e.g., NeRF , mesh, etcetc.) based on the likelihood that a diffusion model defines its 2D projections. To be more specific, these algorithms consist of 3 main components:

1- A 3D parametrization of the scene ϕ{\bm{\phi}}.

2- A differentiable renderer gg to create an image x{\bm{x}} (or its encoded feature) from a desired camera viewpoint vv such that x=g(ϕ,v){\bm{x}}=g({\bm{\phi}},v).

3- A pre-trained 2D diffusion model θ{\bm{\theta}} to obtain a proxy of log⁡p(x∣c,v)\log p({\bm{x}}|{\bm{c}},v) where pp is the 2D data density and c{\bm{c}} is the text prompt.

The 3D generation has been done as solving an optimization problem as follows:

where L\mathcal{L} is a proxy to the negative log-likelihood of the 2D image based on the pre-trained diffusion model.

Remind the noise prediction loss in Equation 3 is a natural choice for L\mathcal{L} as the training objective of the diffusion model, since it is a (weighted) evidence lower bound (ELBO) of the data density :

However, direct optimization of LDiff\mathcal{L}_{\text{Diff}} does not provide realistic samples . Therefore, Score Distillation Sampling (SDS) has been proposed as a modified version of the diffusion loss gradient ∇ϕLDiff\nabla_{{\bm{\phi}}}\mathcal{L}_{\text{Diff}}, which is more robust and more computationally efficient as follows:

where also ϵθ{\bm{\epsilon}}_{\bm{\theta}} has been replaced with ϵ^θ\hat{\bm{\epsilon}}_{\bm{\theta}} to allow text conditioning by using the classifier-free guidance .

Intuitively, this loss perturbs x{\bm{x}} with a random amount of noise corresponding to the timestep tt, and estimates an update direction that follows the score function of the diffusion model to move to a higher-density region.

For the choice of L\mathcal{L}, since the introduction of the seminal work DreamFusion , there have been several proposals . However, since they are similar in core and our method can be applied to all of them, we continue the formulation of the paper by using the Score Distillation Sampling loss presented by DreamFusion.

Since the introduction of 2D diffusion-based 3D generative models, it has been known that they suffer from the Janus (multi-faced) problem . This refers to a phenomenon that the learned 3D scene, instead of presenting the 3D desired output, shows multiple canonical views of an object in different directions. For instance, when the model is asked to generate a 3D sample of a person/animal, the generated object model has multiple faces of the person/animal (which is their canonical view) instead of having their back view.

View-dependent prompting (e.g.e.g., adding back view, side view, or overhead view with respect to the camera position to the main prompt) has been proposed as a remedy but does not fully solve the problem . We believe part of the reason is that 2D Diffusion models fail to be fully conditioned on the view provided by the prompt, as also pointed out by others . For instance, when the model is asked to generate the back view of a peacock, it wrongly produces the front view instead, as the front view has been more prominent in the training data the model has been trained on.

To provide an intuitive mathematical understanding of the Janus problem, we believe one of the reasons is the model fails to be properly conditioned on view vv. More specifically, the proxy of log⁡p(x∣c,v)\log p({\bm{x}}|{\bm{c}},v) does not fully restrict x{\bm{x}} to have zero density on areas that do not represent the viewpoint vv for the scene description yy. The main reason we think this is the case is samples of the density fail to reflect the direction of interest.

2 Perp-Neg to alleviate Janus problem and 2D view conditioning

In this section, we first explain how combining Perp-Neg with a unique prompting technique can enable us to accurately condition the 2D diffusion model on the desired view. Additionally, we will explore how Perp-Neg can be integrated with DreamFusion to address the Janus problem by improving the view faithfulness of the 2D model.

To begin, we demonstrate how to generate a desired statistical view using the improved model. Then, we explain the process for creating interpolations between two views To generate a specific view of an object, we use a combination of positive and negative prompts. We define txtback\textbf{txt}_{\textit{back}}, txtside\textbf{txt}_{\textit{side}}, and txtfront\textbf{txt}_{\textit{front}} as the main text prompts appended by back, side, and front views, respectively. We replace simple prompts containing the view with the following set of positive and negative prompts to generate each view:

where w(⋅)≥0w_{(\cdot)}\geq 0 denotes the weights for the negative prompts. Positive and negative prompts are fed into the Perp-Neg algorithm during each iteration of the diffusion model. We don’t include txtback\textbf{txt}_{\textit{back}} as a negative prompt for the generation of side/front views since most objects’ canonical view is not back. However, if the back view is more prominent for some objects, it should be included as a negative prompt. We also observed increasing the weight of the negative prompt makes the algorithm focus more on avoiding that view, acting as a pose factor.

In this subsection, we will first explain how we interpolate between the side and back views, followed by the interpolation between the front and side views. We distinguish between these two cases because the diffusion model may be biased toward generating front views, and if this assumption is not true, then the formulation needs to be adjusted accordingly.

To interpolate between the side and back views, we use the following embedding as the positive prompt:

where embv\textbf{emb}_{v} is the encoded text for the view vv and rinterr_{\textit{inter}} is the degree of interpolation. And for the negative prompts, we use:

For interpolation between the front and side views, the embedding for the positive would be:

Perp-Neg SDS: We employed interpolation technique in Stable DreamFusion and varied rinterr_{\textit{inter}} based on the related direction of 3D to 2D rendering. To be more specific, we modified the SDS loss 10 as follows:

such that ϵ^θPN(xt;c,v,t)\hat{\bm{\epsilon}}^{\textit{PN}}_{\bm{\theta}}({\bm{x}}_{t};{\bm{c}},v,t) is:

The unconditional term ϵθunc{\bm{\epsilon}}^{unc}_{{\bm{\theta}}} refers to ϵθ(xt,t){\bm{\epsilon}}_{\bm{\theta}}({\bm{x}}_{t},t), and

where c.(v)c_{.}^{(v)} refers to the text embedding of positive/negative at direction vv. And ϵθnegv(i)⊥{\bm{\epsilon}}^{\text{neg}^{(i)\perp}_{v}}_{{\bm{\theta}}} is the perpendical component of ϵθnegv(i){\bm{\epsilon}}^{\text{neg}^{(i)}_{v}}_{{\bm{\theta}}} on ϵθposv{\bm{\epsilon}}^{\text{pos}_{v}}_{{\bm{\theta}}}. And, wvw_{v}’s are representative of the weights of the negative prompts at direction vv.

Further extension: Although we only provide an application of Perp-Neg using SDS loss, we will investigate its application in novel view synthesis , conditional 3D generation , editing and adding texture .

Experiments

In this section, we first conduct experiments on 2D cases to quantitatively demonstrate the importance of using Perp-Neg in the sampling to improve the likelihood of getting the image corresponding to the text query, which provides evidence of why our method surpasses vanilla sampling in the 3D case. Next, we show results in 3D generation.

To understand why Perp-Neg improves the 3D generation quality, we first explore the 2D generation of the requested view to see whether Perp-Neg produces images with fewer artifacts than the vanilla sampling method.

In the first experiment, we fix the random seeds as 0-49 to get 50 images from each text prompt. We carefully select qualified images that align with the requested text based on a series of criteria and report the percentage of accepted samples produced with Stable Diffusion, Compositional Energy-based Model (CEBM), and our Perp-Neg. Below we introduce the details of prompt design and the criteria for accepting qualified samples.

Design of prompts: We design the basic text prompts as: “A [O], [V] view.” Token [O] stands for the objects, such as panda, lion; token [V] stands for view, where we only consider “front”, “back” and “side” in our experiments. For example, we use “A panda, side view” to request the model to generate an image showing the side view of a panda. We aim to test two groups of text prompts that generate the side view and the back view of the objects, which are considered simple and complex cases, respectively. For each group, when using the negative prompts, we use the complementary view or the combination of the other two views, e.g., in the case of using the “side” view in the positive prompt, we use the “front” view in the negative prompt. Both positive and negative prompts follow the basic prompt pattern but respectively adopt positive and negative weight in the fusing stage.

Average success rate: We test each group of prompts using three objects, “panda”, “lion”, and “peacock” and only count photo-realistic generation that matches the text prompt query as a successful generation. For detailed acceptance criteria, please refer to Appendix A.2. On the side view and back view generation, in each group, we adopt 3 combinations of the complementary view into the negative prompts, e.g., for the case that side view is used in the positive prompt, we use front view, back view and both front and back view in the negative prompt. Then we count the averaged percentage of accepted successful generations, summarized in Table 1.

As shown, we can observe the vanilla sampling from Stable Diffusion only has 42.0% in successfully generating requested side-view images. For more difficult cases like generating the back-view images, the success rate is even lowered to 14.6%. By simply using negative prompts without considering the overlap between positive and negative prompts, CEBM fails to generate the desired view and results in a lower success rate compared to Stable Diffusion. Compared to these two types of baseline, Perp-Neg shows the effectiveness of properly using negative prompts and significantly improves the success rate by a large margin. Figure 5 shows a qualitative justification corresponding to Table 1. From the left column, we can observe without using negative prompts Stable Diffusion may generate incorrect views though requested for the back view. Although using negative prompts, CEBM does not consider the overlap between positive and negative prompts, resulting in artifacts or vanish of the content, shown in the middle column. Different than the previous two failed cases, Perp-Neg is able to properly use negative prompts to eliminate the wrong view and preserve the corresponding details of the text prompt query to achieve realistic generation well-aligned to the input text.

On the combination of positive and negative prompts: To explore how to combine negative prompts with positive prompt. We compute the averaged successful generation count across all tested objects and report the averaged count using different positive and negative prompt combinations in Figure 6. From the figure, it is notable that when generating the side-view images, using the back view as the negative prompt is less effective than using the front view or using the combination of both the front view and back view. Similarly, when generating back-view images, using the front view in negative prompts is also less effective, since the model is less likely to generate front-view details when conditioned on the back view, while the side view is more ambiguous to the model. This observation indicates putting ambiguous perspectives in the negative prompt could help the model avoid generating undesired images.

2 Perp-Neg DreamFusion

We integrated Perp-Neg with DreamFusion by using the publicly available replication of DreamFusion that utilizes Stable Diffusion as the pre-trained 2D diffusion model instead of Imagen. We replaced the SDS loss with the one provided in Equation 11. To determine the negative prompt weight functions ff, we used the general form of a shifted exponential decay of the form f.(r)=aexp⁡(−b∗r)+cf_{.}(r)=a\exp(-b*r)+c, where aa, bb, and cc are greater than or equal to zero. We set the parameters of the ff functions separately for each text prompt by generating 2D interpolation samples for 10 different random seeds and then selecting the parameters with the highest accuracy in following the conditioned view. We observed that parameters that allow better interpolation in 2D cases directly relate to the parameters that better help with the Janus problem in 3D. We also noted that if the interpolation between two views remained unchanged while varying the angle for a large range, then the 3D-generated scene was more likely to have a flat geometry from multiple views, resulting in the Janus problem. To overcome this, we perturbed the interpolation factor rr with random noise, calculated the interpolated text embedding and their related negative weights, and thus, the model is less likely to generate identical photos from a range of views.

To evaluate the effectiveness of the Perp-Neg in alleviating the Janus problem, we conducted our experiments using prompts that did not depict circular objects. For each prompt, we utilized the Stable DreamFusion method with and without the Perp-Neg algorithm, running 14 trials for each approach with different seeds. Our results indicate the number of successful outputs generated without a Janus problem when using the Perp-Neg algorithm: “a corgi standing” 2 times, “a westie” 5 times, “a lion” 1 time, “a Lamborghini” 5 times, “a cute pig” 0 times, and “Super Mario” 4 times. In contrast, when we ran the model without the Perp-Neg, it failed to generate any correct output except for “a Lamborghini” 4 times and “Super Mario” 2 times. These findings clearly demonstrate the advantages of utilizing the Perp-Neg in mitigating the Janus problem.

Conclusion

We introduce Perp-Neg, a new algorithm that enables negative prompts to overlap with positive prompts without damaging the main concept. Perp-Neg provides greater flexibility in generating images by enabling users to edit out unwanted concepts from initial generated photos. More importantly, Perp-Neg enhances prompt faithfulness by preventing the 2D diffusion model from producing biased samples from its training data and accurately representing the input prompt. This can be accomplished by feeding to Perp-Neg a sentence describing the model bias as the negative prompt to generate desired solutions. Our paper also demonstrates how Perp-Neg can properly condition the 2D diffusion model to generate views of interest rather than a canonical view. Finally, we integrate Perp-Neg’s robust view conditioning property into SDS-based text to 3D models and show how it alleviates the Janus problem.

References

Supplementary Materials

Appendix A Experiment details

To clarify the difference between with/without negative prompts, and with/without Perp-Neg, we provide an illustrative depiction in Figure 8. Our Perp-Neg does not require additional training or fine-tuning, and is implemented in the sampling pipeline. A detailed implementation in each timestep is shown in Algorithm 1.

For 2D generation, we implement our Perp-Neg into Stable Diffusion v1.4 pipeline. We adopt 50 DDIM steps, and fix the guidance scale as a=7.5a=7.5 and the positive weight wpos=1w_{\text{pos}}=1, for each generation. For negative weight, the value may vary and we find generally in the range [-5, -0.5] can produce satisfactory images. In the 2D view generation experiments, we fix the negative weights to wneg=−1.5w_{\text{neg}}=-1.5 when there is only one negative prompt, and set wneg1 ⁣= ⁣wneg2 ⁣= ⁣−1w_{\text{neg}_{1}}\!=\!w_{\text{neg}_{2}}\!=\!-1 when there are two negative prompts. The results are generated across seeds 0-49. For the interpolation between two views, we normally interpolate rinterr_{inter} with stride 0.25 between 0 and 1 to have 5 images in total. Moreover, for 3D generation, we employ Perp-Neg into Stable Dreamfusion. The results of our baselines are reproduced with their open-source codehttps://github.com/energy-based-model/Compositional-Visual-Generation-with-Composable-Diffusion-Models-PyTorchhttps://github.com/ashawkey/stable-dreamfusion. All experiments are conducted using Pytorch 1.10 on a single NVIDIA-A5000 GPU.

A.2 Criteria for successful view generation count

Here we elaborate on the criteria for the successful generation count in our quantitative experiments in Section 4.1. We reject the image samples with the following criteria and examples are shown in Figure 9:

The images that do not show requested object(s) or view. Note that if the generated image contains multiple objects, and one of them is not positioned in the correct view, the image will still be rejected.

The images show hallucination including counterfactual details, for example, a panda has three ears.

The images have color or texture artifacts that make the images not realistic.

Appendix B Additional experiments

We further conduct case-by-case studies for the previous experiments, where we collect statistics of every tested object per view. We report the averaged acceptance rate across all possible positive and negative combinations in Table 2.

From the table, we find back view is consistently more difficult to generate than the side view. A possible explanation is that the generator may have more information about the front view from the training datasets, and the side view possesses more connection to the front view, while it requires additional knowledge to generate the corresponding back view. In terms of the objects, we find it is less likely to generate faithful images of peacock(s) in the side view, as well as peacock(s) in the back view lion in the back view, as there may have less corresponding training images in the original training datasets.

B.2 Ablation study

We also provide ablation studies on the effects of negative prompt weights to see how the weight affects the usage of negative attribute elimination. We show the results of generations using different negative prompt weights wiw_{i} in Figure 10-12. We can observe for CEBM, the results are consistently not relevant to the requested text content, no matter the negative prompt weights are small or big. For Perp-Neg, we can see with larger weights, e.g., w=−0.1w=-0.1, the generated results are not positioned in the side view. As we decrease the weight, the generated image becomes more relevant to the text. This observation indicates Perp-Neg has better controllability in eliminating the negative attributes in the prompts.

B.3 Additional results

In the following, we provide additional qualitative results of 2D generation and view interpolation. And for 3D generation results, please refer to the provided video in the supplementary file.