ImageDream: Image-Prompt Multi-view Diffusion for 3D Generation

Peng Wang, Yichun Shi

Introduction

In the domain of 3D generation, incorporating images as an additional modality for 3D generation, compared to methods relying solely on text , offers significant advantages, as the common saying, An image is worth a thousand words. Primarily, images convey rich, precise visual information that text might ambiguously describe or entirely omit. For instance, subtle details like textures, colors, and spatial relationships can be directly and unambiguously captured in an image, whereas a text description might struggle to convey the same level of detail comprehensively or might require excessively lengthy descriptions. This visual specificity aids in generating more accurate and detailed 3D models, as the system can directly reference actual visual cues rather than interpret textual descriptions, which can vary greatly in detail and subjectivity. Moreover, using images allows for a more intuitive and direct way for users to communicate their desired outcomes, particularly for those who may find it challenging to articulate their visions textually. This multimodal approach, combining the richness of visual data with the contextual depth of text, leads to a more robust, user-friendly, and efficient 3D generation process, catering to a wider range of creative and practical applications.

Adopting images as an additional modality for 3D object generation, while beneficial, also introduces several challenges, Unlike text, images contain a multitude of features like color, texture, spatial relationships that are more complex to analyze and interpret accurately with a solely encoder like CLIP . In addition, high variant of light, shape or self-occlusion of the object can lead to inaccurate and in-consistent view synthesis, therefore leading blurry or incomplete 3D models.

The complexity of image processing necessitates advanced, computationally intensive algorithms to accurately decode visual information and ensure consistent appearance across multiple views. Researchers have employed various strategies with diffusion models, such as Zero123 , and other recent works , to elevate a 2D object image to a 3D model. However, a limitation of image-only solutions is that, although the synthesized views are visually impressive, the reconstructed models often lack geometric accuracy and detailed textures, particularly in the object’s rear views. This issue primarily stems from significant geometric inconsistencies across the generated or synthesized views. Consequently, during reconstruction, non-matching pixels are averaged in the final 3D model, leading to indistinct textures and smoothed geometry.

Fundamentally, image-conditioned 3D generation represents an optimization problem with more stringent constraints compared to text-conditioned generation. Hence, achieving optimized 3D models with clear details is more challenging, as the optimization process is prone to deviating from the trained distributions due to the limited amount of 3D data. For example, generating a horse based solely on text descriptions may yield detailed models if the training dataset includes a variety of horse styles. However, when an image specifies particular textures, shapes, and fur details, the novel-view texture generation may easily deviate from the trained distributions.

In this paper, we introduce ImageDream to address these challenges. Our approach involves considering a canonical camera coordination across different object instances and designing a multi-level image-prompt controller that can be seamlessly integrated into the existing architecture. Specifically, the canonical camera coordination mandates that the rendered image, under default camera settings (i.e., identity rotation and zero translation), represents the object’s centered front-view. This significantly simplifies the task of mapping variations in the input image to 3D. The multi-level controller offers hierarchical control, guiding the diffusion model from the image input to each architectural block, thereby streamlining the path of information transfer.

As illustrated in Fig.1, ImageDream excels in generating objects with correct geometry from a given image, enabling users to leverage well-developed image generation models for better image-text alignment than purely text-conditioned models like MVDream . Furthermore, ImageDream surpasses existing state-of-the-art (SoTA) zero-shot single image 3D model generators, such as Magic123 , in terms of geometry and texture quality. Our comprehensive evaluation in the experimental section (Sec. 4), which includes both qualitative comparisons through user studies and quantitative analyses, demonstrates ImageDream’s superiority over other SoTA methods.

Related Works

We recognize that 3D generation is a well-established field; this review focuses on significant advancements closely related to our research.

Text-to-3D Generation with Diffusion. The emergence of deep generative models has significantly impacted 3D generation. Early methods targeted the reconstruction of simple objects using multi-view rendered images . The evolution of these techniques, from Generative Adversarial Networks (GANs) to diffusion-based frameworks, marks a notable progression.

Recent 3D diffusion models, specifically for tri-planes and feature grids , have emerged. However, these models often focus on specific objects like faces and ShapeNet objects. Concurrently, there’s growing interest in reconstructing object shapes from monocular image inputs , demonstrating the evolving stability of image generation methodologies. A significant challenge remains in generalizing these models to the extent of their 2D counterparts, likely due to constraints in 3D data size, representation, and architectural design.

Lifting 2D Diffusion for 3D Generation. In light of the limited generalizability of direct 3D generative models, a parallel line of research has explored the elevation of 2D diffusion priors into 3D generation, often integrating with 3D representations like NeRF . A pivotal approach in this area is the score distillation sampling (SDS) introduced by Poole et al., using diffusion priors as score functions to guide 3D representation optimization. Alongside Dreamfusion, works like SJC, which utilize stable-diffusion models , have emerged. Subsequent studies have focused on enhancing 3D representations , refining sampling schedules , and optimizing loss designs . Despite their ability to generate photorealistic objects of various types without 3D data training, these methods struggle with multi-view consistency. Moreover, each 3D model requires individualized optimization through prompt and hyper-parameter adjustments. Notably, MVDream enhances generation robustness by joint training with 2D and 3D datasets, producing satisfactory results with uniform parameters, drawing on multi-view diffusion via SDS. Our work builds upon these concepts, applying them to image-prompt generation and retaining the robustness characteristic of MVDream.

Image-based Novel View Synthesis. Direct synthesis of novel 3D views from single images has also been explored, bypassing traditional reconstruction processes. Watson et al. pioneered diffusion model applications in view synthesis as the pipeline in Sitzmann et al. using the ShapeNet dataset. Subsequent advancements include Zhou et al.’s extension to latent space with an epipolar feature transformer and Chan et al.’s approach to enhance view consistency. Szymanowicz et al. proposed a multi-view reconstructor using unprojected feature grids. A common limitation across these methods is their dependency on specific training data, with no established adaptability to diverse image inputs. Fine-tuning pre-trained image diffusion models on extensive 3D render datasets for novel view synthesis, as proposed by Zero123 , remains constrained by geometric consistency issues. Later works, including SyncDreamer , Consistent 1-to-3 , and Zero123plus , have sought to enhance multi-view consistency through joint diffusion processes, but the reconstruction of geometrically coherent 3D models remains a challenge.

Single Image-conditioned Reconstruction. Recent advances in deriving 3D models from single or few images predominantly leverage NeRF representations. Techniques such as RegNeRF , which uses geometry loss from depth patches, and SinNeRF , RealFusion , and NeuralLift , which combine depth maps or Score Distillation Sampling during NeRF training, represent significant steps forward. Despite their effectiveness, the quality of these generated models remains suboptimal for real-world applications. Magic123 combines single-view and novel-view diffusion networks, achieving impressive texture quality in 3D models. However, our tests reveal limitations in understanding correct object geometry.

We also note recent parallel developments, such as Wonder3D , which incorporate normal diffused outputs into original diffusion models, and DreamCraft3D , which employ a second-stage DreamBooth-like model fine-tuning for enhanced texture modeling. These works, while promising, remain distinct from our contributions.

Methodology

In this section, we first talk about the MVDream pipeline and then describe our method to input the image prompt.

In MVDream, there are two stage for 3D model production. The first stage is training a multi-view diffusion network that produces four orthogonal and consistent multi-view images from a text-prompt given respective camera embedding. In the second stage, a multi-view score distillation sampling (MV-SDS) is adopted to produce a detailed 3D NeRF model.

In the first stage, each block of the multi-view network contains a densely connected 3D attention on the four view images, which allows a strong interaction in learning the correspondence relationship between different views. To train such a network, it adopts a joint training with the rendered dataset from the Objaverse and a larger scale text-to-image (t2i) dataset, LAION5B , to maintain the generalizability of the fine-tuned model. Formally, given text-image dataset X={x,y}\mathcal{X}=\{{\boldsymbol{x}},y\} and a multi-view dataset Xmv={xmv,y,cmv}\mathcal{X}_{mv}=\{{\boldsymbol{x}}_{mv},y,\boldsymbol{c}_{mv}\}, where x{\boldsymbol{x}} is an latent image embedding from VAE , yy is a text embedding from CLIP , and c\boldsymbol{c} is their self-desgined camera embedding, we may formulate the the multi-view (MV) diffusion loss as,

here, x{\boldsymbol{x}} is the noisy latent image generated from a random noise ϵ\epsilon and image latent, the ϵθ\epsilon_{\theta} is the multi-view diffusion (MVDiffusion) model parametrized by θ\theta.

After the model is trained, the MVDiffusion model can be inserted to the DreamFusion pipeline, where the authors adopt a score-distillation sampling (SDS) based on the four generated views. Specifically, in each iteration step, a random 4 orthogonal views are rendered from a NeRF g(ϕ)g(\phi) with a random 4 view camera extrinsic and intrinsic c\boldsymbol{c}. Then, they are encoded to latents xmv{\boldsymbol{x}}^{mv} and inserted to the multi view diffusion network to compute a diffusion loss in the image space which is back propagated to optimize the NeRF parameters. Formally,

Here, x^0mv\hat{{\boldsymbol{x}}}^{mv}_{0} s the denoised MV image at timestep from MVDiffusion. After fusion, MVDream shows significant improvement of object geometry correctness without the Janus issues.

2 Canonical Camera

In the context of MVDream, a critical observation is the diffusion of multi-view images using a global aligned camera coordination. In other words, the image from a default camera (no azimuth rotation) is always the front view of the object. This is done by asking the CLIP image feature of a view best match the ”front view” CLIP text feature embeddings. This alignment facilitates the fusion of diffused images in the fusion step, reducing ambiguity regarding their viewpoints in relation to the provided text prompt.

As emphasized in the introduction, this alignment also reduced the difficulties in learning the accurately reconstructing the geometry of objects. Thereby, in image prompt cases, in contrast to previous image-conditioned approaches like Zero123 , which attempt to recover object 3D geometry based on image camera coordination system, ImageDream adopts canonical/world camera coordination as in MVDream. Our diffusion model aims to regress towards the canonical multiple view image of the object as depicted in the image. This approach is expected to yield superior geometric accuracy compared to systems that utilize relative camera coordination.

Formally, for an image of an object rendered from a random viewpoint with a random camera, denoted as xr{\boldsymbol{x}}_{r}, we create the ImageDream diffusion multi-view (MV) dataset as Xmv={xmv,y,xr,cmv}\mathcal{X}_{mv}=\{{\boldsymbol{x}}_{mv},y,{\boldsymbol{x}}_{r},\boldsymbol{c}_{mv}\}, where cmv\boldsymbol{c}_{mv} is the introduced canonical cameras in MVDream. Then, the rest of diffusion loss is the same as Eqn.(1).

3 Multi-level Controllers

In order to insert the image prompt to control the output MV images, we consider a multi-level strategy. The overall structure of the multi-level controller from an image prompt can be seen in Fig. 4, and we elaborate the details of each component in the following.

Global Controller. In our initial approach, we integrated global CLIP image features into MVDream, akin to how text features are used, by fine-tuning the model’s already well-established training. Recognizing that MVDream is primarily trained on text embeddings, we introduced a multi-layer perceptron (MLP) θg\theta_{g}, functioning as an adaptor similar to IP-Adaptor , following the CLIP image global embedding. This step aims to align image features with text features, ensuring compatibility within the MVDream framework. Specifically, CLIP image encoding encodes image feature to a 1024 vector with a token length of 4, which we named as fg\boldsymbol{f}_{g}. And, θg\theta_{g} further adapts the image feature to be 1024 as the input to cross-attention.

On the MVDiffusion side, inside of an attention layer ll, we add a new set of MLPs, θkg,l\theta_{k_{g},l} and θvg,l\theta_{v_{g},l}, that takes the input the adapted features and output its attention key and value matrix, which then aggregated based on the query feature matrix, ql\boldsymbol{q}_{l}, yielding a corresponding image cross-attention feature hg,l\boldsymbol{h}_{g,l}. Here, a weight λ=1.0\lambda=1.0 is introduced to balance the hidden from text and image, and the final output of layer ll is hl=ht,l+λhg,l\boldsymbol{h}_{l}=\boldsymbol{h}_{t,l}+\lambda\boldsymbol{h}_{g,l}. We refer to decoupled cross-attention in IP-Adaptor for additional details.

To train such a model, we freeze the diffusion model, and only fine-tune {θg,θkg,l,θvg,l}l\{\theta_{g},\theta_{k_{g},l},\theta_{v_{g},l}\}_{l}. We follow the training setting of MVDream by considering both 3D rendered datasets and 2D image datasets together, which will be elaborated in our experimental section.

After the model is tuned, we found the model is able of absorb some informations from the image such as structure of the object etc. As illustrated in Fig. 4 (a), comparing with the input, the diffused output is able to put the pirate hat similar to the image on the bulldog, while some detailed pose and appearance information is lost, which we think is not enough for a good control from the input image.

Local Controller. To enhance control, we try to utilize the hidden feature from the CLIP encoder before its global pooling, which likely contains more detailed structural information. This hidden feature, denoted as fh\boldsymbol{f}_{h}, has a token length of 257257 and a feature dimension of 12801280. A MLP adaptor θh\theta_{h} is introduced to feed fh\boldsymbol{f}_{h} into the diffusion network’s cross-attention module, with θkh,l\theta_{k_{h},l} and θvh,l\theta_{v_{h},l} forming the key and values matrix. These parameters, {θkh,l,θvh,l,θh}l\{\theta_{k_{h},l},\theta_{v_{h},l},\theta_{h}\}_{l}, are then jointly trained as learnable elements similar to the global controller. Post-training, we observed that the results were overly sensitive to image tokens, leading to overexposed and unrealistic images, especially with higher class free guidance (CFG) settings , as shown in Fig. 4(b).

To mitigate this, we implemented a resampling module θr\theta_{r}, following the approach of IP-Adaptor, reducing the hidden token count from 257 to 16, resulting in a more balanced local image feature fr\boldsymbol{f}_{r}. The corresponding local controller parameters are {θr,θkr,l,θvr,l}l\{\theta_{r},\theta_{k_{r},l},\theta_{v_{r},l}\}_{l}. As Fig. 4(c) illustrates, after this resampling, the diffused images more realistic, even at higher CFG levels. From the generated images, it’s evident that the model captures the overall layout and object shape, but also struggles with finer identity details like object skin texture.

Pixel Controller. To optimally integrate object appearance texture, we propose embedding the image prompt pixel latent x{\boldsymbol{x}} across all attention layers in ImageDream. Specifically, MVDream employs a 3D dense self-attention mechanism with a shape of (bz,4,c,hl,wl)(bz,4,c,h_{l},w_{l}) across four views within a transformer layer. In contrast, ImageDream introduces an additional frame by concatenating the input image, resulting in a feature shape of (bz,5,c,hl,wl)(bz,5,c,h_{l},w_{l}). This enables similar 3D self-attention processes between the four-view images and the input image.

During the training of our diffusion network, we refrain from adding noise to the latent from the input image prompt, ensuring the network clearly captures the image information. Additionally, to differentiate the input image features and avoid confusion, we assign an all-zero vector to the camera embedding of the input image. Given that the pixel controller is integrated into the multi-view diffusion without extra parameters, we fine-tune all feature parameters in unison, adopting the same training regime as the global/local controllers but with a learning rate reduced by a factor of ten. This approach preserves the original feature representations more effectively. Post-training, as depicted in Fig 4(d), the generated multi-view images not only ethically maintain the appearance from the input image but also uphold the multi-view consistency characteristic of MVDream, resulting in satisfactory 3D model fusion.

Finally, our multi-level controller is a combined one with local and pixel, since we think the global one do not have too much additional information. There might be potential queries regarding the necessity of a pixel controller, given that IP-Adaptor, relying solely on CLIP features, can capture extensive image texture details: we posit that while IP-Adaptor is effective for modifying objects in the same view as the input, decoding the same view is comparatively simpler. Multi-view diffusion, however, presents a more complex challenge. Decoding from a highly compressed CLIP feature could necessitate prolonged training on larger datasets. Therefore, at this stage, we find the pixel controller significantly beneficial for rapidly training a robust multi-view diffusion model.

4 Image-Prompt Sore Distillation

Implementing the image-prompt multi-view diffusion network in ImageDream follows the multi-view score distillation framework of MVDream (Sec. 3.1), with the addition of an image prompt as an input to the diffusion network. However, we need to condier few key differences in NeRF optimization to achieve accurate results.

Background Alignment. During SDS optimization, the NeRF-rendered image includes a randomly colored background to differentiate the interior and exterior of the 3D object. This random background, when input into the diffusion network alongside the object, can conflict with the background from the image prompt, leading to floating artifacts in the generated NeRF model, as shown in Fig. 5(a). To resolve this, we adjusted the image-prompt background to match the rendered background color from NeRF, successfully eliminating these artifacts.

Camera Alignment. Our diffusion network tends to generate multi-view images mirroring the camera parameters (e.g., elevation, field of view (FoV)) of the input image prompt, parameters which remain unknown during NeRF rendering. Randomly sampling parameters for rendering, as done in MVDream, can result in images incongruent with the image prompt’s rendering settings, affecting the geometry of detailed image structures. To mitigate this, we narrowed the parameter sampling range from MVDream’s ,, for camera FoV and elevation to andand, respectively, a range more typical for a generated user photos. This adjustment significantly improved the geometric accuracy of the 3D objects, as demonstrated in Fig. 5(b).

We acknowledge this solution’s limitations; when the image prompt’s camera parameters greatly differ from our selected range in canonical camera setting, the resulting 3D object shape may be unpredictable. Future improvements could include a camera parameter estimation module or increased randomness in the image prompt rendering during diffusion training, to better synchronize the settings between NeRF rendering and diffusion.

Experiments

In this section, we detail the experimental setup for ImageDream, designed to enable replication of our model. We will release both the model and code following this submission.

Implementation Details. Adhering to the dataset configuration of MVDream (Sec. 3.3), we used a combined dataset from Objaverse for 3D multi-view rendering and a 2D image dataset for training controllers (Sec. 3.3). For image prompts in the 3D dataset, we randomly selected one of the 16 front-side views, with azimuth angles ranging from $degrees,outofthetotal32circleviews.Forthe2Ddataset,weusedtheinputimageastheimageprompt.Arandomdropoutrateof0.1wassetfortheimagepromptduringtraining,replacingdroppedpromptswitharandomuni−coloredimage.Forallexperiments,i.e.withglobalcontroller,localcontrollerandlocalpluspixelcontrollers,wetrainedfordegrees, out of the total 32 circle views. For the 2D dataset, we used the input image as the image prompt. A random dropout rate of 0.1 was set for the image prompt during training, replacing dropped prompts with a random uni-colored image. For all experiments, i.e. with global controller, local controller and local plus pixel controllers, we trained for60Kstepswithabatchsizeof256andagradientaccumulationof2,usingtheAdamWoptimizer.Themodelisinitializedfromstablediffusion(2.1)checkpointofMVDream.Thelearningratewassetto1e−4,exceptforthemodelwithpixelcontroller,whereitwasreducedto1e−5.TestimagepromptswereresizedtoK steps with a batch size of 256 and a gradient accumulation of 2, using the AdamW optimizer. The model is initialized from stable diffusion (2.1) checkpoint of MVDream. The learning rate was set to 1e-4, except for the model with pixel controller, where it was reduced to 1e-5. Test image prompts were resized to256\times 256,andwesetthediffusionCFGto5.0.Thetrainingtakes, and we set the diffusion CFG to 5.0. The training takes\sim2dayswith2 days with8A100.ForNeRFoptimization,wefollowedMVDream’sconfigurationbutintroducedathree−stageoptimizationatresolutionsA100. For NeRF optimization, we followed MVDream’s configuration but introduced a three-stage optimization at resolutionswhichswitchedat[which switched at [5K,K,10K]steps,settingthecameradistancebetweenK] steps, setting the camera distance between[0.6,0.85]$ for better NeRF model coverage. The NeRF training is about 1hr with A100.

Test Dataset. Our primary focus was evaluating ImageDream outside the Objaverse distribution to ensure its practical applicability. We selected 39 well-curated prompts from MVDream, covering a diverse range of objects with relatively complex geometries and appearances, surpassing datasets like ShapeNet or CO3D . Using SDXL , we generated multiple images from each prompt, selecting ones with aesthetically pleasing objects. The backgrounds of these images were then removed, and the objects re-centered, akin to the approach used in Zero123.

In our evaluation, we compared ImageDream’s performance against several SoTA baselines, including Zero123-XL (trained on 10x larger data than ours), Magic123 , and SyncDreamer . The criteria for comparison were geometry quality and similarity to the image prompt (IP). ’Geometry quality’ refers to the generated 3D asset’s conformance to common sense in terms of shape and minimal artifacts, while ’similarity to IP’ assesses the resemblance of the results to the input image. We executed all baseline tests using default configurations as implemented in threestudiohttps://github.com/threestudio-project/threestudio.

Qualitative Evaluation. Lacking ground truth for the test image prompts, we conducted a real user study to evaluate the quality of the generated 3D models. Participants were briefed on our evaluation standards and asked to choose their preferred model based on these criteria. The experiment was double-blind, with participants shown 3D assets generated by different methods without identifying labels. The comparison results, depicted in Fig. 6, show that ImageDream, both with (ImageDream-P) and without (ImageDream-G) the pixel controller, significantly outperformed other baselines. ImageDream-P was particularly favored, while ImageDream-G also received a positive preference rate. SyncDreamer was omitted from the figure due to its NeuS results having a 0%\% preference rate.

Fig. 7 presents a representative case comparing results from the diffusion models and the final NeRF model. Systems like Magic123 and Zero123, which rely on single-view diffusion with relative camera embedding, often produce incorrect geometry, as illustrated by their inability to accurately represent the span of horse body. In contrast, ImageDream, through its unique design, effectively resolves this issue, resulting in more satisfactory models (more results are list in webpage).

Numerical Evaluation. To thoroughly assess image quality at various stages of our generation pipeline, we employed the Inception Score (IS) and CLIP scores using text-prompt and image-prompt, respectively. The IS evaluates image quality, while CLIP scores assess text-image and image-image alignment. However, since IS traditionally evaluates both image quality and diversity within a set, and our prompt quantity is limited, the diversity aspect makes the score less reliable. Therefore, we modified the IS by omitting its diversity evaluation, replacing the mean distribution with a uniform distribution. Specifically, we set qiq_{i} in IS to be 1/N1/N, making the IS of an image ∑ipilog⁡(Npi)\sum_{i}p_{i}\log(Np_{i}), where NN is the inception class count and pip_{i} is the predicted probability for the ithi_{th} class. We denote this modified metric as Quality-only IS (QIS). For the CLIP score, we calculated the mean score between each generated view and the provided text-prompt or image-prompt.

In Tab.1, we present comparative results. SD-XL, reflecting the score of test images, achieved the highest QIS and CLIP scores. MVDream, listed as a benchmark for final 3D model quality, shows improved synthesized image quality after 3D fusion due to multi-view consistency. In contrast, Zero123 and Zero123-XL experienced a drop in image quality post-3D fusion due to diffusion inconsistency. Magic123 enhanced the CLIP score over Zero123 by integrating a joint diffusion model. SyncDreamer’s quality declined as it diffuses only 16 fixed views, complicating reconstruction. In ImageDream, we evaluated three models for ablation: one with a global controller, another with a local controller, and the last incorporating both local and pixel controllers (Sec.3.3). ImageDream maintained high image quality in both diffusion and post-3D fusion stages. The local controller, in particular, provided better image CLIP scores post-fusion, thanks to richer image feature representations. The pixel controller model excelled in image CLIP scores during both stages. Notably, ImageDream-pixel ranked second in other scores, with Zero123-XL using a significantly larger dataset (Objaverse-XL ).

However, these scores don’t fully encapsulate important aspects like multi-view consistency and geometric correctness. For instance, as Fig. 7 demonstrates, zero123-XL, despite having high IS due to easy image classification, showed poorer consistency. Thus, while these scores offer some reliability when consistency is high, future research should aim to develop more comprehensive metrics that accurately capture geometric correctness to better compare different generation algorithms.

2 Limitations

While the model incorporating the pixel controller achieves the best scores, we observed certain trade-offs, particularly when the image constraints are overly stringent. For instance, in cases like the small facial details of a full-body avatar (shown in Fig.8), the model struggles to capture these nuances, whereas global control might recover the face based on the text prompt. To address this, as outlined in Sec.3.4, the pixel controller model needs to better estimate image intrinsic and extrinsic properties, or a better balance tuning inside of multi-level controllers. This may be solved by exploring the use of larger models, such as SDXL , which be our future work.

Conclusion

We introduce ImageDream, an advanced image-prompt 3D generation model utilizing multi-view diffusion. This model innovatively applies canonical camera coordination and multi-level image-prompt controllers, enhancing control and addressing geometric inaccuracies seen in prior methods. Future improvements could focus on increasing randomness in image-prompts during training to further reduce texture blurriness in the generated models. These steps are expected to further advance the capabilities and applications of ImageDream in 3D model generation.

Acknowledgements and Ethics Statement

We thank our 3D group members of Kejie Li, and our intern Zeyuan Chen in joint meeting discussion and setup baselines of SyncDreamer, which help complete this paper. In addition, note that the models proposed in this paper aims to facilitate the 3D generation task that is widely demanded in industry for ethical purpose. It could be potentially applied to unwanted scenarios such as generating violent and sexual content by third-party fine-tuning. Built upon the Stable Diffusion model , it might also inherit the biases and limitations to generate unwanted results. Therefore, we believe that the images or models synthesized using our approach should be carefully examined and be presented as synthetic. Such generative models may also have the potential to displace creative workers via automation. That being said, these tools may also enable growth and improve accessibility for the creative industry.

References