SODA: Bottleneck Diffusion Models for Representation Learning

Drew A. Hudson, Daniel Zoran, Mateusz Malinowski, Andrew K. Lampinen, Andrew Jaegle, James L. McClelland, Loic Matthey, Felix Hill, Alexander Lerchner

Introduction

What I cannot create, I do not understand.

Synthesis, the ability to create, is considered among the highest manifestations of learning . As opposed to passive analysis of a text or an image, conceiving them out of thin air involves profound understanding of the underlying factors and intricate generative processes that give rise to the final product . Indeed, learning to write in a new language is often more challenging than reading it. Figuring out the solution to a math problem is fundamentally harder than verifying it . And just as the chef learns more about the culinary arts than the diner to prepare a tasty meal, and the novelist knows more about narrative structures than the reader to tell a good story, the artist better grasps perspective and composition to craft a breathtaking masterpiece.

Analogously, in AI, the recent years have witnessed remarkable progress at the generative domain, with large-scale diffusion modeling proving to be a powerful and flexible technique that can create vivid imagery of astonishing realism and incredible detail. And yet, while the vast majority of research harnesses these models for the straightforward goal of synthesis or editing alone , only little attention has been given to their representational capacity , leaving this promising direction rather unexplored. Surely, models that can weave from scratch such rich depictions of high fidelity, likely learn much along the way about the underlying properties, processes, and components that make up the resulting pictures. How then can we leverage this untapped potential of diffusion models for the purpose of representation learning, and extract the knowledge they acquire for the benefit of downstream tasks?

Motivated to achieve this aim, we present SODA, a self-supervised diffusion model, designed for both perception and synthesis. It couples an image encoder with the classic diffusion decoder , both trained in tandem for novel view generation – a task we choose to employ here, not only for its own sake, but as a self-supervised objective. The encoder transforms an input view into a concise latent representation, which then guides the denoising of an output view, by modulating the decoder’s activations.

This setup introduces a desirable information bottleneck between the encoder and the decoder , that in contrast to the typical diffusion framework, equips our model with an explicit and interpretable visual latent space. As our experiments confirm, its advantages are twofold: it both encourages the emergence of disentangled and informative representations that capture image key properties and semantics, which thus can be applied to downstream tasks, and further provides effective means to control and manipulate the produced outputs, for the gain of image editing and synthesis. We further devise and integrate multiple new ideas into the network architecture and training procedure: layer modulation, modified classifier-free guidance, and an inverted noise schedule, so to maximize its representation skills.

We demonstrate our model’s strengths and versatility by evaluating it along a series of classification, reconstruction and synthesis tasks, spanning an extensive collection of datasets that covers both the simulated and real-world kinds. SODA possesses strong representation skills, attaining high performance in linear-probing experiments over the ImageNet dataset among others. Moreover, it excels at the task of few-shot novel view generation, and can flexibly synthesize images either conditionally or unconditionally, as indicated by metrics of fidelity, consistency and diversity. Finally, we inspect the model’s emergent latent space and discover its disentangled nature, which offers controllability over the semantic traits of the images it produces, as validated both qualitatively and quantitatively.

Overall, SODA integrates together three research ideas that we seek to establish and promote: First, diffusion models are not only adept at image generation, but are also capable of learning strong representations. Second, novel view synthesis can serve as a powerful self-supervised objective for model pre-training. And third, the compactness of the latent space, which could be reached by constricting the bottleneck between the encoder and the denoiser, plays a pivotal role in enhancing the latent representations’ quality, informativeness and interpretability.

Related Work

Diffusion. The advent of diffusion models has lately marked a breakthrough in the field of visual synthesis. Originally inspired by theories of thermodynamics , it approaches generative modeling by following a reversible and iterative denoising process, the forward direction of which slowly erodes the structure within the data distribution, while the backward direction is gradually restoring it. Since its early inception back in 2015 , tremendous strides have been made in the quality and diversity of the created outputs, thanks to innovations of the framework’s training and sampling techniques . Consequently, diffusion models have been widely adopted for numerous tasks and modalities , synthesizing images, videos, audio and text , and even advancing planning and drug discovery , effectively becoming one of the leading paradigms for generative modeling nowadays.

But while most literature highlights its generative feats, only a handful of works have studied diffusion modeling’s representational capacity, mainly repurposing pre-trained text-to-image models for classification , segmentation , or multimodal reasoning . The reliance on such models makes it unclear whether the downstream capabilities arise from the diffusion approach itself, or are actually attributable to the exceptionally large scales, long training and voluminous captioned data, which, essentially, provides rich and textual semantic supervision. To address this shortcoming, we focus here instead on the fully-unsupervised regime, and train our model from scratch on standardized benchmarks, seeking to asses the value and potential of diffusion-based representations derived from images alone.

Visual Encoding. Closer to our work is DRL , that extends early research on denoising auto-encoders , and conditions a denoiser on an encoded clean version of its own target. It is mainly explored from a theoretical perspective, along with preliminary results on MNIST and CIFAR-10. DiffAE follows up, integrating style modulation into the encoder , while InfoDiffusion regularizes it with mutual-information loss. Our approach builds upon this line of research, but instead of auto encoding the same image, we generate novel views. Notably, we discover that this, in turn, remarkably enhances the model’s representation skills, as evidenced by substantial gains in downstream performance. We further couple this idea with multiple technical innovations, pertaining both architecture and optimization, geared to realize the representational capabilities of diffusion models to their fullest. And in contrast to prior works, we provide an extensive empirical study of diffusion-based representation learning, encompassing a broad suite of datasets over multiple different tasks.

Hybrid Models. A couple of partially related works are unCLIP and Latent Diffusion Models , both of which utilize a frozen pre-trained encoder (CLIP and VQGAN respectively) to cast images onto a compressed latent space over which a diffusion model can operate. Consequently, we note that the latent representations used in both these approaches are in fact not derived by diffusion itself, but rather through either contrastive or adversarial pre-training. As such, they differ fundamentally from our study, which aims to explore the effectiveness of diffusion-based pre-training as a means for representation learning.

Downstream Tasks. For each of the tasks we explore – classification, disentanglement, reconstruction, and novel view synthesis – we compare SODA to the leading prior works. These include models such as SimCLR, DINO, and MAE for linear-probe classification , NeRF-based approaches for novel view generation , and classic variational models for the task of disentanglement . Whereas these techniques are designed for particular objectives or depend on domain-specific assumptions, SODA exhibits a greater degree of versatility, as it tackles representational and generative tasks alike.

Approach

SODA is a self-supervised diffusion model that learns a bidirectional mapping between images and latents. It consists of an image encoder E(x′)=z\mathcal{E}(\bm{x^{\prime}})=\bm{z} that casts an input view x′\bm{x^{\prime}} into a low-dimensional latent z\bm{z}, which is then used to guide the synthesis of a novel output view x\bm{x}, that relates to the input x′\bm{x^{\prime}} (Figure 2). Concretely, x\bm{x} is produced through a diffusion process that is conditioned on the encoding z\bm{z} via feature modulation . This design equips SODA with an explicit and compact latent space, which not only offers ample control over the generative process, but can also be leveraged for downstream perception tasks (Section 4)Our model is named after the soda drink. Indeed, the fizzing in soda bottles is an everyday example of the diffusion phenomena..

We first present an overview of the model (Section 3.1, Figure 2), followed by an in-depth discussion of each of its core components: the encoder’s architectural design (Section 3.2), the mechanisms involved in the synthesis of novel views (Section 3.3), and the optimization techniques we develop to cultivate strong and meaningful representations (Section 3.4).

As a denoising diffusion model , SODA is formally defined by a pair of forward and backward Markov chains, iteratively transforming a sample xT\bm{x}_{T} from the normal distribution into the target one (x0\bm{x}_{0}) and vice versa. Each forward step tt erodes xt\bm{x}_{t} by adding low Gaussian noise ϵt\bm{\epsilon}_{t} according to a fixed variance schedule αt\alpha_{t}. Meanwhile, the respective backward step performs image denoising, and aims to estimate ϵt\bm{\epsilon}_{t} in order to recover xt−1\bm{x}_{t-1} from its successor xt\bm{x}_{t}. It is carried out by a decoder D\mathcal{D}, implemented as a convolutional UNet (with 2m+12m+1 activation layers hih_{i}).

To tackle the denoising challenge, we assist the decoder by conditioning it on a latent vector z\bm{z}, which guides its operation by modulating the activations h\bm{h} at each of its layers through adaptive group normalization that controls their scale and bias: zsGroupNorm(h)+zb\bm{z}_{s}\text{GroupNorm}(\bm{h})+\bm{z}_{b} where (zs,zb)(\bm{z}_{s},\bm{z}_{b}) are linear projections of z\bm{z}, applied evenly across the activation grid. See Appendix B for the architectural details of the decoder, as well as closed-form equations of the diffusion forward and backward steps. The latent z\bm{z} is created by the newly introduced image encoder E\mathcal{E}, discussed next.

2 The Encoder

The latent representation z\bm{z} is at the model’s core, serving as the communication channel between the encoder E\mathcal{E} and the denoising decoder D\mathcal{D}, while guiding the latter through the diffusion process. It is derived by a ResNet encoder E(x′)=z\mathcal{E}(\bm{x^{\prime}})=\bm{z}, representing a clean source view x′\bm{x^{\prime}} that semantically or visually relates to the target view x\bm{x} (Section 3.3).

The driving motivation behind this idea stems from the denoising task’s inherent under-determination: Seeking to fulfill it to the best of its ability, the decoder will leverage any pertinent knowledge or useful clue that could inform it of the missing content to fill in. This in turn incentivizes the encoder to distill into z\bm{z} the most striking and prominent commonalities between the source and target views. Since we further constrict the encoding z\bm{z} into a low-dimensional space, and use it to guide the decoder via global feature modulation, it yields a fruitful combination that encourages z\bm{z} to specifically capture the image’s high-level semantics, while delegating the reconstruction of localized and high-frequency details to the denoiser itself.

From that perspective, learning a latent z\bm{z} that supports image denoising, rather than pure reconstruction from scratch, as in most auto-encoders , liberates our encoder from the need to compress all information about the image into the representation, and let it instead focus on the image’s most distinctive and descriptive qualities.

To enhance the model’s latent space disentanglement, we introduce two intertwined mechanisms of modulation and masking. In layer modulation, we partition the latent vector z\bm{z} into m+1m+1 sections – half the number of layers in the decoder D\mathcal{D}. We use each zi\bm{z}_{i} to modulate the respective pair or layers (hi,h2m−i)(h_{i},h_{2m-i}), thereby promoting specialization among the latent sub-vectors. Just like light rays refracted through a prism, they are encouraged to capture visual traits at different levels of granularity, from coarser to finer, so to guide the decoder’s operation through the layers.

To improve the localization and reduce the correlations among the mm sub-vectors, we present layer masking – a layer-wise generalization of classifier-free guidance . During training, we zero out a random susbset of z1:m\bm{z}_{1:m}, effectively performing layer-wise guidance dropout, that mitigates the decoder’s reliance on sub-vector dependencies, allowing them to decouple and specialize independently. At sampling, we extrapolate the model’s output in the conditional direction: ϵθ(xt∣z)−ϵθ(xt∣0)\bm{\epsilon}_{\theta}(\bm{x}_{t}{\mid}\bm{z})-\bm{\epsilon}_{\theta}(\bm{x}_{t}{\mid}\bm{0}). This endows SODA with finer control over the generative process, and opens the door for image editing and style mixing , as we can selectively condition the decoder on some levels of granularity, like structural or positional aspects, while giving it free rein to unconditionally vary other ones, such as lighting, texture, or color palette (see supplementary figures).

3 Novel View Generation

We loosely consider views to be any set of images that hold some relation among each other, such as visual or semantic (Figure 2): they can be various augmentations or distortions of an original image, as is commonly explored in the contrastive learning literature , they can show a 3D object from different poses and perspectives , or they can simply share the same semantic category with one another.

Cardinality. We permit the trivial singular case where all views are identical, which then turns the model into an auto-encoder. Conversely, we can extend the conditioning and create a novel view based on a set of kk input views, instead of just a single one. We map each input xix^{i} to its latent with a shared encoder E\mathcal{E}, and aggregate the resulting latents z1:k\bm{z}^{1:k} into a single vector z\bm{z}, either by taking their mean, or by processing them through a shallow transformer.

Perspective. The model can incorporate richer forms of conditional information, such as the camera perspective associated with each view: Specifically, for experiments over 3D datasets like ShapeNet (Section 4.2.2), we concatenate a grid of ray positions and directions r=(o,d)\bm{r}=(\bm{o},\bm{d}), embedded with sinusoidal positional encoding , to the linearly-mapped RGB channels of the source and denoised views, x′\bm{x^{\prime}} and xt\bm{x}_{t}. This allows us to conditionally generate novel views that match the requested pose and orientation. See supplementary for illustrations and implementation details.

Guidance. We extend classifier-free guidance for novel view synthesis, and instead of masking just the latent z\bm{z}, we randomly and independently mask either the latent or the pose information r\bm{r}. As our empirical findings suggest, this idea not only hones the model’s generative skills, but further enables conditioning on partial information, allowing SODA to either unconditionally conceive novel objects at a requested pose, or alternatively, generate arbitrary novel views of given objects at the absence of source or target positional information, based on an image only (Appendix G).

Cross Attention. The technique of layer modulation offers the encoding z\bm{z} with global control over each of the decoder’s layers, by evenly setting their scales and biases across the grid. We further study alternative mechanisms, and explore the integration of cross attention, so to support spatial modulation. Instead of layer-wise modulation (which we use by default), we partition the latent z\bm{z} into nn sub-vectors, and perform cross attention between the sub-vectors set and the decoders’ activations, akin to the word-based attention common in text-to-image generation. We find that cross attention aids the model at 3D novel view synthesis, while layer modulation performs better for image editing, reconstruction, and representation learning.

4 Training & Sampling

Noise Schedule. We train the model with the standard MSE objective , but introduce a new noise schedule to better fit the representation learning task: Indeed, diffusion models commonly set the variance of the additive noise term ϵ\epsilon to follow either a cosine , sigmoid or linear decay schedules, prioritizing noise levels that are close to the margins, either of the high or low ends (Figure 3). Those schedules have been found useful for image synthesis.

However, from a representation learning perspective, denoising images with overly high or low noise levels fails to provide effective training signal for the model to learn from: Too little noise does not present the denoiser with a challenging enough task, thereby diminishing the encoder’s necessity. Meanwhile, too heavy noise puts excess pressure on the latent z\bm{z} to fully capture every pixel-level detail of the image, turning denoising into mere reconstruction. We thus incorporate a new inverted noise schedule, that promotes medium noise levels in lieu of the extremes, which proves highly conducive to representation quality (Section 4.1).

Additional Settings. Two modifications we find beneficial for representation learning are: (1) adding low Gaussian noise to the encoder’s input images; (2) optionally setting the encoder to have a higher learning rate than the decoder, so to positively impact their learning dynamics by allowing the encoder to adapt faster as it guides the decoder in the denoising task. While the model is robust to the selection of the learning-rate ratio, tuning it could improve downstream results. Once trained, we use DDPM for sampling.

Experiments

We evaluate SODA through a suite of quantitative and qualitative experiments, demonstrating its strong representation skills and generative capabilities over 12 different datasets grouped into 4 tasks: We begin with linear-probe classification (Section 4.1), showing the model’s utility for downstream perception. We proceed to image reconstruction and few-shot novel view synthesis (Section 4.2), illustrating its ability to envision 3D objects from new unseen perspectives. We then explore the model’s disentanglement & controllability (Section 4.3), as substantiated by comparative analysis and latent-space interpolations.

In the supplementary and our website website (soda-diffusion.github.io), we provide additional samples, visualizations and animations, and provide further details about our evaluation procedures, discussing the datasets, metrics, baselines, and implementation details (Appendices C-F). We conclude with ablation studies (Appendix G), that empirically validate the contribution of the model’s components and design choices. Taken altogether, the evaluation offers solid evidence for the efficacy, robustness and versatility of our approach.

We assess the quality of SODA’s learned representations through linear-probe analysis , over ImageNet and CelebA, which complement one another: the former calls for fine-grained clustering into 1000 possible categories, while the latter involves rich identification of diverse semantic attributes. We train our model in a self-supervised fashion: using RandAugment for ImageNet, and Gaussian data augmentation for CelebA. We then fit a linear classifier over the latent vectors z\bm{z} that predicts the category or attributes, and measure the resulting performance. We note that training diffusion models for representation learning is computationally efficient, since iterative sampling is necessary for generative purposes only.

As shown in Table 1, SODA reaches 72.24% accuracy (top1) on the ImageNet1K linear-probe classification task, outshining competing generative approaches such as MAE, BEIT and iGPT , and significantly reducing the gap with discriminative and constrastive approaches like DINO, SwAV, and SimCLR . Meanwhile, for CelebA, our model attains the strongest results (72.7% F1) compared to competing approaches (Supp Table 6), eclipsing even the language-supervised CLIP embeddings (71.1% F1).

SODA proves remarkably robust to the choice of data augmentation, as it performs strongly regardless of the selected strategy, seeing only a minor decrease of 3.1% when switching the heavier RandAugment for a lighter crop+flip augmentation. This stands in stark contrast to the high sensitivity of contrastive methods to data augmentations, with e.g. BYOL and SimCLR suffering from major drops of 13.6% and 27.5% respectively when light augmentation is applied (crop+flip), and other approaches relying on particular schemes such as MultiCrop among others.

The newly introduced image encoder plays an instrumental role in the model’s downstream performance, enabling a 3x boost over features obtained from an unconditional diffusion model The baseline obtains image encodings by pooling the features of the best-performing denoiser layer (a middle one) over a lightly-noised image., which for ImageNet scores 24.49% only. A significant improvement further arises from the use of novel view synthesis as a self-supervised representation learning objective: Indeed, maintaining a distinction between the model’s source and target views yields a 17.12% increase for ImageNet, compared to when they match (i.e. auto-encoding), even though the same data augmentation is applied in both cases (Figure 6)For auto-encoding, we sample a new augmentation at every training step, but use it both as the source and as the target..

Other contributors include the compact bottleneck and feature modulation, which respectively raise accuracy by 38.04% and 15.51% for ImageNet, and 12.74% and 11.25% F1 for CelebA (Supp Table 11). Designing a new inverted noise schedule that favors medium noise levels in lieu of the extremes likewise strengthen the model’s representation capacity, eliciting a 10.25% increase over ImageNet.

2 Visual Synthesis

Next, we analyze the model’s generative skills, evaluating its ability to faithfully reconstruct an image x\bm{x} from its latent encoding z\bm{z} (Section 4.2.1)To evaluate our model at the task of reconstruction, we keep the source the and the target views the same during training x′=x0\bm{x^{\prime}}=\bm{x}_{0}., and generate novel views of 3D objects from requested camera perspectives, given one or more conditional source views (Section 4.2.2). We evaluate the targets and predictions’ similarity along multiple dimensions: pixel-wise (PSNR ), structural (SSIM ), perceptual (FID ) which accounts for sharpness and realism, and semantic (LPIPS ) (Appendix E.2).

As indicated by Table 2, SODA produces excellent reconstructions, surpassing competing approaches like VQGAN , StyleGAN2 , DALL-E and unCLIP (DALL-E2) , especially in terms of structural and semantic similarity (SSIM & LPIPS). Visually speaking, our samples are sharper and crispier than DALL-E’s, while being more accurate than StyleGAN2 and unCLIP inversions, perhaps due to their lack of a trainable encoder (see supplementary examples). The results are significant given the order-of-magnitude lower dimensionality of the latents z\bm{z} from which we restore the images: 2K for SODA versus 65-524K for VQGAN and DALL-EThe overall latent dimension is 16×16×25616{\times}16{\times}256 – 32×32×51232{\times}32{\times}512., illustrating here an advantage of continuous representations over discrete codebooks.

2.2 Novel View Generation

For the task of few-shot novel view synthesis, we focus on the 3D regime and look into 3 datasets that span both synthetic object renderings and real-world scans of household items (Google Scanned Objects , custom ShapeNet , and NMR ). We compare our model to geometry-free and -aware approaches such as PixelNeRF , Scene Representation Transformer (SRT) , and Palette . We condition the models on 1-9 source views, and test them on held-out validation objects that do not appear in training.

SODA consistently beats the competing approaches across the 3 datasets and for different numbers of source views, as indicated by FID, SSIM and LPIPS (Table 3 and Supplementary Figure 7). It reaches the largest gains along LPIPS and FID, producing significantly sharper images that better match the source views both structurally and semantically. We observe that settings of 1-3 source views benefit the most from our model, where for the single-source case, it improves FID scores by an order-of-magnitude and often almost halves the LPIPS scores. For the GSO dataset, as we increase the source views number, we score a little lower on the pixel-wise PSNR than PixelNeRF, perhaps due to the probabilistic nature of our approach. Yet, in terms of computational efficiency, contrary to the slow and heavy rendering of geometry-aware methods, SODA maintains strong performance with as little as 20 sampling steps.

Figure 5 and the supplementary animations feature objects synthesized from various perspectives, showcasing the viewpoint consistency SODA achieves. For multiple sources, we find that our proposed transformer-based view aggregation (Section 3.3) surpasses the stochastic conditioning technique of 3DiM (Figure 6). Our approach further outperforms the denoiser-only Palette diffusion model, which fits translational tasks that closely follow the source layout, like colorization or super-resolution, but struggles at structural transformations, corroborating the need for our dedicated image encoder.

3 Disentanglement & Controllability

The concept of disentanglement has been a recurring theme in representation learning research over the years . While formal definitions may vary , a common aim lies in the discovery of abstract and meaningful latent representations that linearly align with the natural axes of variation. Disentanglement could enhance the encodings’ interpretability, and, in the context of generative modeling, support greater controllability. In the following, we inspect the model’s latent space, and analyze it quantitatively and qualitatively along the dimensions of disentanglement, controllability and informativeness.

Latent Interpolations. We begin by visualizing latent interpolations for our model, linearly traversing the latent space from one vector z1\bm{z}_{1} to another z2\bm{z}_{2} (Figure 1 and the supplementary figures & animations). We observe smooth variations over traits of texture and structure. Notable in particular are image categories that seamlessly morph from one to another (e.g. from a tiger to a cat to a dog to a wolf), while for CelebA, we see gradual transformations ranging from broad shifts of pose and orientation to finer transitions of hair, facial features and expressions.

Attribute Manipulation. We go beyond interpolations and identify meaningful latent directions that correspond to individual axes of variation. To infer them, we explore two techniques: supervised, by normalizing the linear-probe’s weight matrix, and unsupervised, through PCA decomposition (details at Appendix E.3). Figure 4 and the supplementary figures show perturbations along the discovered directions, which influence various semantic properties: from age and gender in CelebA, to tone, clarity and lighting conditions in LSUN, to object’s dimensions, thickness and material in ShapeNet and GSO. Indeed, we see that these manipulations are disentangled: one attribute is altered while the others are mostly kept intact, demonstrating the well-behaved nature of SODA’s emergent latent space and the quality and strength of its learned representations. Indeed, we emphasize that both SODA’s training and the following PCA-based discovery of interpretable latent directions are fully-unsupervised, derived from images only.

Layer Modulation. We investigate the effect of layer modulation and masking (Section 3.2.1), blocking the encoder’s guidance from select decoder’s layers and inspecting the impact on the produced outputs. As illustrated in the supplementary, it allows for selective modification of input images at different levels of granularity, so to preserve certain factors while unconditionally regenerating other ones. We can thus resample a new color palette while retaining shape and structure, or change the background while preserving the subject’s identity. We find that layer masking improves the model’s robustness to these forms of partial conditioning. Meanwhile, the role played by the initial Gaussian noise map xT\bm{x_{T}} is closely linked to the chosen augmentation scheme: it controls fine stochastic subtleties like fur or freckles when SODA is trained to reconstruct the source image, and could conversely shape the underlying layout when heavier data augmentations are applied.

3.2 Quantitative Evaluation

To quantitatively bolster the findings above, we analyze our approach with DCI , which measures representations along Disentanglement, Completeness and Informativeness by assessing the degree of 1-to-1 correspondence between latent and ground-truth factors of variation (Appendix E.3). We evaluate models over multiple semantically-annotated datasets, ranging from the diagnostic SmallNORB and MPI3D to the realistic CUB and CelebA .

As Table 4 and supplementary Tables 6 and 7 show, SODA outshines both variational and adversarial approaches, improving Disentanglement by 27.2-58.3% and Completeness by 5.0-23.8% across 4 different datasets, with the sole exception of the synthetic 3DShapes, for which both SODA and most variational methods attain excellent scores. For Informativeness, results are mostly comparable, with SODA taking the lead for some datasets, while StyleGAN or DIP-VAE improving scores for others. Our experiments further validate the contribution of layer modulation and masking, respectively yielding 3.2-13.5% and 2.5-8.9% increases in latent-space Disentanglement, and 3.8% and 1.7% mean increase in Completeness. Visually, SODA’s samples are significantly sharper than the variational ones, and it achieves remarkable boosts in realism (FID) and semantic similarity (LPIPS).

Conclusion

We introduced SODA, a self-supervised diffusion model, designed for both perception and synthesis. It re-purposes the task of novel view generation as a training objective for representation learning. By conditioning a denoiser on an image encoder, and imposing an information bottleneck between the two, SODA learns strong semantic representations that enable downstream classification, as well as reconstruction, editing and synthesis. While we focused on single-object images, as in LSUN, ShapeNet, or ImageNet, we believe that exploring the applicability of our approach to dynamic compositional scenes is a promising direction for future research. We hope our work will help bridging the gap between novel view synthesis and self-supervised learning, two flourishing topics that are often pursued independently, and bring us one step closer to unlocking the potential of generative models in general and diffusion models in particular to advance the representational frontier.

References

Supplementary Material

Appendix A Overview

In the following, we discuss additional analysis of our approach, and provide further description of the model structure, implementation details, and evaluation procedures. Appendix B offers an overview of diffusion models’ preliminaries and equations. In Appendix C, we then specify the chosen hyperparameters, training techniques, and sampling methods. Appendices D, E and F respectively review the datasets, metrics, and baselines we consider in this study. Finally, in Appendix G, we present ablation and variation studies that assess the contribution of each of our design choices, complementing the principal ones explored in the main paper.

We plan very soon to add to the supplementary and our website (soda-diffusion.github.io) a variety of animations and visualizations of outputs generated by the model over different datasets, spanning image reconstructions, viewpoint traversals, latent interpolations, unsupervised attribute discovery and manipulation, demonstration of style and content (or structure) separation, qualitative impact of layer masking and variation of the initial noise map for different training data augmentation schemes, and samples conditioned on partial information.

Appendix B Model Overview & Diffusion Preliminaries

where αˉt=∏s=1tαs\bar{\alpha}_{t}=\prod\nolimits_{s=1}^{t}{\alpha_{s}} is the product of the variances up to step tt, and σt2\sigma_{t}^{2} is either a fixed or learned variance term. To train the model, we can readily obtain xt\bm{x}_{t} with the closed-form computation (where ϵ∼N(0,I)\bm{\epsilon}\sim\mathcal{N}(\bm{0},\bm{I})):

and couple it with the simplified re-weighted MSE training objective (where ϵθ\bm{\epsilon}_{\theta} is estimated by the model):

In terms of the architecture, our model consists of an image encoder E\mathcal{E} (ResNet or ViT), and a denoising decoder D\mathcal{D} that follows the classic structural design of prior literature , featuring a UNet implemented as a stack of residual, convolutional, and either downsampling or upsampling layers (in the encoding and decoding modules of the UNet respectively), that are further linked by symmetric skip connections. The decoder D\mathcal{D} notably integrates Adaptive Group Normalization layers throughout, allowing z\bm{z} and tt to modulate the decoder’s activations of each layer h\bm{h}, by scaling and shifting them channel-wise:

where (ts,tb)(\bm{t}_{s},\bm{t}_{b}) and (zs,zb)(\bm{z}_{s},\bm{z}_{b}) are both obtained by linear projections, the former of a sinusoidal timestep embedding of tt , and the latter of the latent representation z\bm{z} created by the image encoder E\mathcal{E}.

Appendix C Implementation Details

Architecture. See Table 12 for our chosen hyperparameters. In terms of the training objective, optimization scheme and empirical configuration, we adopt most of the common settings of recent works , and specifically use the Adam optimizer , gradient accumulation, and exponential moving average for the model’s weights; for the ResNet encoder : variant v2 ResNet , Xavier initialization , ReLU non-linearity, dropPath , and mean pooling; and for the UNet decoder: truncated normal initialization (JAX default), GeLU non-linearity , 2 \sqrt{2}\, rescaling of residual connections, BigGAN re-sampling order , and self attention in the decoder’s low-resolution layers (8-32).

Training. For each dataset, we train the model until convergence, as measured by lack of improvement over a set number of training steps along a validation metric of choice (either downstream accuracy or SSIM). For sampling, we use discrete-time DDPM , classifier-free guidance and 1000 diffusion timestemps, practically strided into 75-250 steps . We implement SODA in JAX , and run our experiments either on NVIDIA Tesla V100s or TPUs (v2).

Positional Encoding. We employ sinusoidal positional encoding to represent both timesteps and, in the case of pose-conditional view synthesis, spatial coordinates, either xy grids for the 2D case or camera rays’ origins and directions for 3D, normalized to a range of $.Incontrasttotheoriginalencodingschemeusedtorepresentdiscretewordpositions,wefurtherscaletheargumentsof. In contrast to the original encoding scheme used to represent discrete word positions, we further scale the arguments of\sinandand\cosbyafactorofby a factor of2{\pi}s(with(withs$ being a hyperparameter), so to increase the distinction among the positional encodings (Figure 10).

Pose Conditioning. Throughout the paper, we experiment with several different flavors of the novel view synthesis task: either generating a view conditionally, matching a 3D pose or 2D coordinates, or alternatively, in a pose-unconditional fashion: where given a source view, the model is asked to generate arbitrary novel views at perspectives of its choice). For the conditional case, we represent each perspective by a H×WH{\times}W 2D grid – of (x0,y0)×(x1,y1)(x_{0},y_{0})\times(x_{1},y_{1}) in the 2D case, and ray positions and directions in the 3D case – embedded by sinusoidal positional encoding and concatenated to linearly-mapped RGB channels of the corresponding view, after the first layer of the encoder and the denoiser respectively. In Appendix G, we compare different ways to represent the rays, such as through normalization, by casting them on a plane or a sphere, or by summing up their positions and directions.

Learning Rates. For the ImageNet dataset, we maintain a different learning rate between the encoder and the denoiser, at a ratio of lrElrD>1\frac{\text{lr}_{\mathcal{E}}}{\text{lr}_{\mathcal{D}}}>1. We practically implement it by following the idea of learning rate equalization , scaling down the initialized weights of the encoder by a factor of kk (by scaling down the standard deviation of the initialization distribution), and then having the network itself scale them back up by kk, effectively scaling the encoder’s gradients by kk. While the model is robust to the selection of the learning rate ratio, we find that a ratio of 2 yields optimal downstream results (Appendix G).

Appendix D Datasets, Preprocessing & Augmentations

Throughout this work, we evaluate models over various datasets grouped into multiple tasks, as summarized by Table 10 and through the textual description below:

Representation Learning & Reconstruction: Each image in the following datasets is associated with a category label (or for CelebA, with multiple attribute annotations).

Imagenet1K : includes diverse images of objects among 1,000 categories of e.g. animals, instruments, furniture and food items.

CelebA-HQ : features face images, annotated with 40 binary semantic properties like age, gender, or hair color; used also for quantitative disentanglement analysis.

LSUN : partitioned into multiple categories of objects (like cars, cats and horses) and scenes (e.g. bedrooms and churches); See Table 10 for full list.

Animal Faces-HQ (AFHQ) : covers various breeds of cats, dogs and wildlife.

Oxford Flowers 102 : features diverse flowers from the United Kingdom.

Novel View Synthesis: Each image in the following datasets is associated with the camera perspective it was captured from, expressed as a grid of ray positions and directions r=(o,d)\bm{r}=(\bm{o},\bm{d}).

NMR : consists of ShapeNet objects’ renderings at 24 fixed views, evenly spaced around a surrounding ring with constant radius and altitude; images of 64×6464{\times}64 resolution. We use the SoftRas data split .

ShapeNet: our custom ShapeNet renderings dataset, featuring 120 views randomly sampled from an upper hemisphere, with random azimuth ϕ\phi, altitude θ\theta, and radius r ∈ [rmin,rmax]r\,{\in}\,[r_{\text{min}},r_{\text{max}}]; 256×256256{\times}256 resolution; created by the Blender-based Kubric library .

Google Scanned Objects (GSO) : includes scans of real-world household items, which we render with Blender following the same protocol described above.

Disentanglement (Quantitative): Each image in the following datasets is associated with discrete semantic attribute annotations.

SmallNORB : contains toy images belonging to 5 categories like animals and vehicles, captured from various camera perspectives and lighting conditions.

3DShapes : includes images of a centered object among varied combinations of shape, color, size and orientation (4 shapes, 8 scales, 15 orientations, and 10 possible colors for the object, wall and floor).

MPI3D : includes 4 splits of either synthetic or real objects, hold by a robotic arm, with different discrete attributes (4-6 shapes, 4-6 colors, 2 sizes, 3 background colors, and 3×40×40{3}\times{40}\times{40} camera perspectives).

Caltech-UCSD Birds (CUB-200-2011) : contains images of various bird species, annotated with 312 binary semantic properties.

D.2 Data Preprocessing

Resolution. We resize all images for training and evaluation to a source resolution of 256×256256{\times}256, inputted into the encoder E\mathcal{E}, and target resolution 128×128128{\times}128, produced by the decoder D\mathcal{D}, with the exception of CUB and ImageNet: for the former, we center-crop and pad each image based on its associated bird’s bounding box; for the latter, we first resize the target images to 256×256256{\times}256, and then center-crop them to 224×224224{\times}224, matching prior literature .

We keep the model’s output resolution as 128 since according to diffusion models’ practices, higher-resolution images are commonly produced through cascading , where a core module first generates images of resolution 64 or 128, and these are these are subsequently post-processed by an independent super-resolution module, rather than being created as high-resolution directly. Indeed, this technique has been shown to improve the overall sample quality, and could readily fit with our approach as well.

Normalization. We normalize the input images fed into the encoder E\mathcal{E} based on ImageNet mean and variance statistics , while linearly scaling the target images of the denoising decoder D\mathcal{D} to the range $$, following the standard procedures.

Data Splits. For each dataset, we either use the default splits, or if not provided, split them into 80% training, 10% validation and 10% testing. Data is shuffled at training time. We note that for all the multi-view datasets: NMR, ShapeNet, GSO, and smallNorb, we intentionally keep all the views of each object exclusively grouped within one of the splits, and consequently, all the objects used for evaluation are not included in the training set.

D.3 Data Augmentation

We study several augmentation schemes, applied for different tasks and datasets: by default, we use random resize cropping, horizontal flipping and optionally RandAugment data augmentation on both the source and target views x′\bm{x^{\prime}} and x\bm{x} (encoded and denoised respectively). Specifically, at every training step, we randomly augment each view, at the rates specified in Table 12. To train the subsequent downstream classifier, we perform cropping and flipping only, and finally, at evaluation time, perform only center-cropping, following the standard linear probing protocols of prior self-supervision learning works . When training the diffusion model, we also find it conducive to add low Gaussian noise to the encoded source view, similarly to the noise added to the denoised target view.

Meanwhile, for multi-view 3D datasets such as NMR, GSO and ShapeNet, we do not apply data augmentations, and instead, randomly sample one view as the source and another as the target, further supplied by their respective camera perspectives (Appendix C). In this case, we allow for conditioning on multiple source views, and conduct experiments over k ∈ k\,{\in}\, sources. Lastly, to illustrate the ability of SODA to learn useful representations even without relying on data augmentation, we perform ablations on datasets used as is, forgoing augmentations of any kind.

Appendix E Evaluation & Metrics

We explore SODA for multiple types of tasks and purposes: downstream linear-probe classification and disentanglement analysis for assessing the quality of the learned representations, as well as image reconstruction and novel view synthesis for evaluating the model’s generative capabilities. These skills are measured both through qualitative inspection of the latent space, with visualizations that demonstrate its impact on the model’s outputs (including in particular latent interpolations and unsupervised attribute discovery), as well as through an assortment of metrics that quantify each of the capabilities as discussed below.

In Section 4.1, we analyze the model’s learned latent representations by measuring their predictive performance on a downstream classification task. Following the common evaluation protocol , we first train our model on a collection of images, and then fit a linear classifier that considers the latent encodings produced by the model and use them to predict each respective image’s category or semantic attributes. The classifier is either trained on the frozen representations z\bm{z} subsequently to the training of the diffusion model, or alternatively, trained with it in tandem by blocking the gradient flow between the two networks – we find that both approaches achieve similar results.

When training the classifier, we refrain from applying weight decay, and adhere to either light augmentation of cropping and flipping for ImageNet or no augmentation in other cases. The latents z\bm{z} are normalized before being fed to the classifier, concretely, by processing them with an un-parameterized batch normalization , which only tracks mean and variance statistics and lacks the follow-up affine transformation. After normalizing the latents, we use 0.1 dropout for regularization, and for ImageNet, apply label smoothing of 0.1. Since the annotated datasets we explore all have discrete labels, we use softmax cross entropy to train the classifier, and report its performance along metrics such as F1 for binary attributes, and top1 accuracy for other ones.

E.2 Image synthesis

In Section 4.2.1, we analyze the capacity of SODA to both reconstruct an input view and generate novel views. Given a target image xx, we assess the quality of a synthesized output x^\hat{x} through multiple complementary metrics that range from visual to semantic similarity:

Peak Signal-too-Noise Ratio (PSNR) ↑\uparrow (measured in dB) : is directly derived from the mean MSE between xx and x^\hat{x}, and it thereby measures pixel-wise similarity. It may rate a blurry estimation as highly consistent with the target, as long as they match well with each other on average.

Structural Similarity Index Measure (SSIM) ↑\uparrow (ranges between $$) : compares images along three perceptual factors: luminance, contrast and structure, and is thus better correlated with the Human Visual System (HVS).

Learned Perceptual Image Patch Similarity (LPIPS) ↓\downarrow(often normalized to be in $$) : computes the distance between the target and synthesized images in the feature space of a supervised pre-trained network, such as VGG , and therefore serves as an indicator for semantic similarity.

Fréchet inception distance (FID) score ↓\downarrow (is ≥0{\geq}0) : quantifies realism and sharpness of the generated images by comparing their distribution to that of the target ground-truth images. It concretely achieves it by considering the mean and variance of each, in a latent feature space, e.g. of the Inception model . When assessing unconditionally-generated images, the FID score further expresses their diversity, but in the case of conditional synthesis, either as reconstructions or with pose conditioning, it mainly reflects their fidelity, sharpness and lack of distortions (also known as R-FID in this context).

For fair comparison, we compute these metrics over all approaches using the same metrics’ implementations, and specifically, casting the images to the range of forPSNR,for PSNR, for FID and LPIPS, and using a uniform kernel to calculate SSIM scores.

E.3 Disentanglement

In Section 4.3, we examine the latent space of our model and assess its degree of disentanglement and controllability through quantitative and qualitative evaluation methods:

DCI metrics (Disentanglement, Completeness & Informativeness) ↑\uparrow (at a range of [0,100%][0,100\%]) : measures the 1:1 alignment between the latent representation z\bm{z} and the natural (ground-truth) factors of variation c\bm{c}. Disentanglement reflects the extant to which each latent variable z^j\hat{z}_{j} (the jj’th axis of the vector z\bm{z}) corresponds to a unique natural factor cjc_{j}. Completeness inversely measures the extant to which each natural factor cjc_{j} is captured by a single latent variable z^j\hat{z}_{j}. Finally, Informativeness indicates the predictability of the natural factors cc from the latent encoding z\bm{z}. These metrics are derived from the normalized importance matrix of a learned classifier and its performance, where the classifier is based on either gradient boosting or Lasso (we use the former). Our implementation of these metrics closely follows Locatello et al. .

Latent Interpolation: we randomly pick two images x1,x2\bm{x}_{1},\bm{x}_{2} from each dataset, encode them to obtain z1,z2\bm{z}_{1},\bm{z}_{2}, and then decode back the latents along a linear segment that connects between the endpoints: z1+(z2−z1)⋅t\bm{z}_{1}+(\bm{z}_{2}-\bm{z}_{1}){\cdot}t for t ∈ t\,{\in}\,, which results in a visualization of the latent traversal.

Principal Component Analysis (PCA) : we encode a sample set of NN images (1,000-10,000), and perform PCA decomposition over the obtained latents zi=1N\bm{z}_{i=1}^{N}, which yields the latent directions sj\bm{s}_{j} of the greatest variation. We then traverse the latent space along the discovered directions: z+sjt\bm{z}+\bm{s}_{j}t for t ∈ [−λj,λj]t\,{\in}\,[-\sqrt{\lambda_{j}},\sqrt{\lambda_{j}}] where λj\lambda_{j} is the respective eigenvalue and λj\sqrt{\lambda_{j}} is the standard deviation along direction sj\bm{s}_{j}. Doing so allows us to visualize the impact of these latent directions on the model’s generated images, and indeed, we find they strongly correlate with semantically-meaningful manipulations.

Thanks to layer modulation (Section 3.2.1), we can further perform PCA over chosen sub-vectors of z\bm{z} that are responsible for guiding decoder’s layers of interest. This enables the discovery of latent directions that control particular levels of granularity, from low-frequency structural aspects to high-frequency factors like texture and color, enhancing the model’s overall controllability.

Classifier-based Attribute Manipulation: For datasets with binary attributes annotations, such as CelebA and CUB, we can produce similar visualizations to the ones described above by examining the weight matrix’s rows of the linear probes we train for the downstream classification experiments (Section 4.1). Indeed, these probes are trained to capture the latent directions that correspond to the presence or absence of the semantic attribute annotations that accompany the datasets we study. The key difference between the PCA-based approach and this technique is that the former is unsupervised while the latter is not.

Appendix F Baselines

For each of the tasks we explore, we compare our model to the respective leading approaches, as well as to additional ablated baselines that we design. Here, we list and review all the baseline methods we compare to.

First, we implement multiple baselines and ablated models within our diffusion codebase, and report their performance across the range of tasks:

EncDec: a vanilla encoder-decoder x^=D(E(x′))\bm{\hat{x}}=\mathcal{D}(\mathcal{E}(\bm{x^{\prime}})), sharing the same encoder and decoder architectures as SODA (for D\mathcal{D}, the decoding module of the UNet), but being trained to generate the target image x\bm{x} from scratch, with no denoising. We train this model both with and without (x=x′\bm{x}=\bm{x^{\prime}}) data augmentation.

Palette: a diffusion model that instead of having a dedicated encoder E\mathcal{E}, concatenates the source image x′\bm{x^{\prime}} to the denoised image xt\bm{x}_{t} directly, and inputs both of them to a UNet denoiser D(xt,x′)\mathcal{D}(\bm{x}_{t},\bm{x^{\prime}}) (also known as Image-to-Image diffusion model ).

unCLIP (Dall-E2) : a diffusion model that relies on a frozen pretrained CLIP as the encoder E\mathcal{E}. We emphasize that we do not refer here to the already trained Dall-E2 model, but rather to its architecture, and so we train its denoiser (in a comparable size to our model) from scratch along with the frozen pre-trained CLIP encoder, for each dataset of interest.

w/o bottleneck: an ablation of SODA with no bottleneck, which rather encodes the input image x′\bm{x^{\prime}} into a 2D feature grid zw×h\bm{z}^{w{\times}h}, with no global pooling, and conditions the denoising on it through cross-attention (similarly to text-to-image diffusion models ).

w/o modulation: an ablation of SODA that broadcasts and concatenates the latent z\bm{z} to linearly-mapped RGB channels of the denoised image xt\bm{x_{t}}, instead of applying modulation through adaptive group normalization (also called a spatial broadcast decoder ).

F.2 Linear-Probe Classification

For downstream classification, we compare our model to a diverse array of leading self-supervised learning approaches: generative methods like MAE , BEIT and iGPT split each image into a grid of tokens or patches, mask some patches and predict them back from the unmasked ones, oftentimes using a transformer backbone.

Meanwhile, discriminative approaches leverage contrastive learning (as in SimCLR ), clustering techniques (as in SwAV ), and distillation (as in DINO and BYOL ) to derive visual representations. At the core of these methods is a strong reliance on rich data augmentations, which are essentially the driving force that allows the to perform unsupervised clustering.

Consequently, contrastive learning approaches operate well at tasks that involve identification of an image’s category, as is the case for ImageNet, but may struggle to capture finer traits that are altered by the augmentations. The semantic properties they may or may not encode into the learned representations heavily depend on the particularities of the data augmentation scheme they employ, since they are basically encouraged to form a latent space that is invariant to the augmentation applied, instead casting different augmentations into similar representations.

Contrary to these two kinds of approaches, both of which are unsuitable for high-quality image generation, SODA stands out being able to both encode input images into meaningful latents, and also synthesize back crisp output images, conditionally and unconditionally. It learns compact and disentangled representations, which contrast with the large, potentially discrete, 2D grids learned by alternative approaches, and as demonstrated in Section 4.1, is robust to the chosen data augmentation scheme, operating well even in its absence.

Our comparison to the approaches discussed in this subsection relies on the performance reports in their respective publications over the ImageNet1K dataset, with the exception of the crop+flip accuracy for SwAV and DINO’s for which we retrain the models.

F.3 Image Reconstruction

We examine the performance of varied models for the task of image reconstruction: Dall-E and VQGAN employ a discrete variational auto-encoder , which casts input images into 2D token grids, based on a trainable codebook. These approaches then couple the auto-encoder with a prior-distribution model, to enable unconditional image synthesis. However, for our purposes (image reconstruction), we consider the auto-encoder module only.

The adversarial StyleGAN model can also be used for image reconstruction, by applying optimization-based inversion techniques to infer back latents from images. Given an image xx, they leverage gradient descent to reverse engineer the latent zz that gives rise to an output x^\hat{x} that is as close as possible to the image xx while still staying on the model’s learned manifold. While these techniques tend to produce samples that share semantic properties with the source images, they oftentimes fail to reconstruct them faithfully. Finally, we compare our model to the diffusion-based DiffAE , which, in contrast to our study, focuses on auto-encoding only, and can be viewed as a predecessor of our approach, as discussed in Section 2.

We assess the reconstruction capabilities of the approaches described in this subsection by evaluating a sample set of images produced by their associated public pre-trained checkpointed models.

F.4 Novel View Synthesis

For novel view synthesis of 3D objects, we compare SODA to a collection of geometry-free and -aware approaches designed for few-shot settings: PixelNeRF learns to translate a small number of source views into a neural radiance field, and then use volumetric rendering techniques to generate new ones. NeRF-VAE extends this idea by leveraging amortized variational inference to learn probablistic neural scene representations. In contrast to these specialized methods, designed specifically for 3D environments, SODA proves considerably more versatile, successfully addressing a broader spectrum of tasks and datasets.

As an alternative to differentiable rendering, geometry-free approaches often use attention mechanisms to directly transform source views into targets: Scene Representation Transformer (SRT) parametrizes scenes with the computationally lighter and faster Light-Field formulation , and synthesize output views from new perspectives by directly attending to the input views’ encodings. The diffusion-based 3DiM goes further and makes extensive use of cross-attention throughout all of its network’s layers so to directly map sources to targets. In contrast to these approaches, we intentionally introduce a bottleneck into our model that induces a meaningful and compact latent space. This, in turn, offers much tighter control over the model’s generative process, opening the door for both semantic manipulation of given scenes, as well as unconditional synthesis of new ones – two new capabilities that are out of these prior works’ reach.

We evaluate the methods described in this subsection either using the authors’ official implementations (for NeRF-VAE), or with our own re-implementations (for PixelNeRF and SRT), matching the originally reported performance.

F.5 Disentanglement

In terms of disentanglement, we analyze SODA over a suite of semantically-annotated datasets, and compare it with a series of variational approaches , which are traditionally known for encouraging the formation of disentangled representations: β\beta-VAE , re-weights the KL regularization term to constrain the latents’ capacity; AnnealedVAE slowly relaxes the encoder-decoder bottleneck so to foster gradual learning; FactorVAE and β\beta-TCVAE encourage factorization of the latent distribution by reducing the correlations among the axes; DIP-VAE (variants I and II) penalizes the mismatch between the prior and the posterior, so to similarly encourage factorization within the latter. We evaluate these methods using the official disentanglement-lib TensorFlow repository , while modifying the backbone encoder and decoder architectures to match the ones used in SODA, for better comparability.

Appendix G Ablation Studies

To gain better insight into the relative contributions our design decisions make, we conduct thorough ablation and variation studies for each of the model’s components, inspecting the (1) feature modulation used to propagate information between the encoder and the denoiser, (2) data augmentation strategies for the source and target views, encoding and conditioning schemes of (3) positional information for our 3D multi-view experiments, (4) sampling configurations of the denoising process and its classifier-free guidance, and finally, (5) the encoder and denoiser’s respective sizes, dimensions and learning rates.

This study joins ablations presented through the main paper (Sections 4.1 and 4.3.2) that attest to the strengths and benefits of the model’s core aspects and key innovations, like bottleneck compactness (Section 3.2), layer modulation (Section 3.2.1), redesigned noise schedule (Section 3.4), and incorporation of novel view synthesis as a self-supervised training objective (Section 3.3).

We explore multiple modulation variants and examine how they fare in terms of generative skills and downstream performance (Tables 9, 11 and 7). As the results suggest, modulation-based conditioning proves considerably more effective than alternative mechanism such as input concatenation [xt,x′][\bm{x}_{t},\bm{x^{\prime}}] (Palette ) or spatial broadcasting (w/o modulation) , with respective deltas of 15.5% and 12.3% at classification over ImageNet (top1) and CelebA (F1), and 0.32 (out of 1.0) mean SSIM improvement at novel view synthesis. Layer modulation proves beneficial too, enhancing disentanglement scores, with up to 13.9% improvement, and generative capabilities, with 0.13 increase in SSIM and halving of LPIPS for ImageNet reconstructions.

We further assess ways to integrate the guidance of the timestep tt and the latent z\bm{z}, and as an alternative to our two-stage guidance approach, where the denoiser’s activations h\bm{h} are modulated first by tt and subsequently by z\bm{z}, we map them instead to a single pair (ws,wb)(\bm{w}_{s},\bm{w}_{b}) either through summation or concatenation (sum/concat mod.) which is then used to modulate the activations: AdaGN(h,t,z)=wsGroupNorm(h)+wb\text{AdaGN}(\bm{h},t,\bm{z})={\bm{w}_{s}}\text{\text{GroupNorm}}(\bm{h})+{\bm{w}_{b}} However, our two-step strategy proves stronger than this variant. Likewise, scaling the features multiplicatively with zs\bm{z}_{s}, as opposed to adding a bias term zb\bm{z_{b}} only, leads to small improvements across different datasets.

G.2 Data Augmentation

We examine the impact of data augmentations on the model’s performance, and analyze variations of the augmentation method itself as well as the inputs it is applied to (Figure 6). At training, our model receives two inputs: a clean view x′\bm{x^{\prime}} processed by the encoder E\mathcal{E}, and a noisy view xt\bm{x}_{t} denoised by the decoder D(xt∣x′)\mathcal{D}({\bm{x}_{t}}{\mid}\bm{x^{\prime}}), aiming to recover x=x0\bm{x}=\bm{x}_{0}. With the exception of native multi-view datasets (e.g. ShapeNet), we create the source and target views x′\bm{x^{\prime}} and xt\bm{x}_{t} by applying random data augmentations at each training step on the original image x\bm{x} (from the dataset).

We test the impact of applying augmentations either just to the source view, just to the target view, to both, or to none of them. We observe that augmenting the source is more critical than the target in terms of its influence on downstream classification performance, and that the model still achieves 55.1 when the source and the target views remain equal. Moreover, we find it valuable to add low Gaussian noise to the source views read by the encoder, yielding 1.3% improvement in ImageNet classification accuracy and improving LPIPS scores relatively by 33%.

G.3 Pose Conditioning

We compare different encoding schemes of the camera perspectives for the 3D novel view synthesis task (Table 8). Given a camera pose p\bm{p}, we can use a closed-form calculation to derive a 2D grid of rays r=(o,d)\bm{r}=(\bm{o},\bm{d}) of dimension H×W×6H{\times}W{\times}6 with origins o\bm{o} and directions d\bm{d}. We can then represent each ray through concatenation: [o,d][\bm{o},\bm{d}] (concat), as commonly done in prior works , or instead, express them with a parametric sum: o + sd⋅d\bm{o}\,+\,s_{d}{\cdot}\bm{d}, where sds_{d} is a scaling factor that can be chosen in different ways: either normalizing d\bm{d} to a length of 1, casting it onto the image plane, or, as we propose, on a sphere that centers at the object, i.e. the origin (Normalized, Plane and Sphere). We can further describe the rays either using Polar or Cartesian coordinates, embedded with sinusoidal positional encoding as explained above (Appendix C). We compare these alternatives, and find that casting the rays on a sphere performs most effectively, and that for this case, Polar coordinates outperform the Cartesian ones.

We further experiment with representing the camera pose as a single vector p\bm{p} that captures its position and direction in Polar coordinates, either considering pt−ps\bm{p}_{t}-\bm{p}_{s}, the relative camera transformation from the source to the target, or concatenating the two absolute viewpoints [ps,pt][\bm{p}_{s},\bm{p}_{t}]. We then encode the information with sinusoidal positional encoding ), and use the resulting vector to guide the denoiser’s operation through feature modulation, similarly to the latent z\bm{z}. However, the ablations show that integrating the camera perspective by concatenating a 2D grid of rays surpasses both the modulation-based pose conditioning as well as a hybrid alternative that simultaneously uses both techniques.

Finally, we experiment with different masking techniques as part of the classifier-free guidance (Section 3.3), either randomly masking the latent representation z\bm{z} that encodes the source image view, masking the pose information (namely, the rays 2D grid r\bm{r}), or independently masking both. We interestingly note that the ideal masking vary for different datasets: while masking of both the latent and pose improves performance for the real-world Google Scanned Objects, it reduces the performance for ShapeNet. Qualitatively, masking both the pose r\bm{r} and the latent z\bm{z} enhances the model’s generative flexibility, allowing it to synthesize either novel objects at requested camera perspectives, or arbitrary novel views even at the absence of source or target’s pose information (supplementary figures will be added very soon).

G.4 Sampling Configuration

We vary the guidance strength gg and timesteps striding ll (i.e. number of timesteps used at sampling) and analyze their impact on the generated images’ quality along different metrics (Figure 9). The model is robust to variation in both settings, with optimal values commonly achieved at g=2g=2 and l=150l=150 (considering different metrics and datasets). Classifier-free guidance consistently yields higher-quality images than unguided sampling, while too strong guidance (like ≥5{\geq}5) results in a slight reduction in scores and potential visual artifacts.

As per the number of sampling steps, while PSNR and LPIPS scores tend to remain constant, we interestingly observe an inverse correlation between the process length’s influence on FID vs. SSIM, the former reflecting sharpness and fidelity while the latter capturing similarity to the target: Sampling images over more steps tends to improve their realism, but may simultaneously induce subtle variations, as the samples begin to slightly move away from the mean estimated target. As aforementioned, we find that l=150l=150 offers a favorable balance between these two qualities.