Lumiere: A Space-Time Diffusion Model for Video Generation

Omer Bar-Tal, Hila Chefer, Omer Tov, Charles Herrmann, Roni Paiss, Shiran Zada, Ariel Ephrat, Junhwa Hur, Guanghui Liu, Amit Raj, Yuanzhen Li, Michael Rubinstein, Tomer Michaeli, Oliver Wang, Deqing Sun, Tali Dekel, Inbar Mosseri

Introduction

Generative models for images have seen tremendous progress in recent years. State-of-the-art text-to-image (T2I) diffusion models are now capable of synthesizing high-resolution photo-realistic images that adhere to complex text prompts (Saharia et al., 2022b; Ramesh et al., 2022; Rombach et al., 2022), and allow a wide range of image editing capabilities (Po et al., 2023) and other downstream uses. However, training large-scale text-to-video (T2V) foundation models remains an open challenge due to the added complexities that motion introduces. Not only are we sensitive to errors in modeling natural motion, but the added temporal data dimension introduces significant challenges in terms of memory and compute requirements, as well as the scale of the required training data to learn this more complex distribution. As a result, while T2V models are rapidly improving, existing models are still restricted in terms of video duration, overall visual quality, and the degree of realistic motion that they can generate.

A prevalent approach among existing T2V models is to adopt a cascaded design in which a base model generates distant keyframes, and subsequent temporal super-resolution (TSR) models generate the missing data between the keyframes in non-overlapping segments. While memory efficient, the ability to generate globally coherent motion using temporal cascades is inherently restricted for the following reasons: (i) The base model generates an aggressively sub-sampled set of keyframes, in which fast motion becomes temporally aliased and thus ambiguous. (ii) TSR modules are constrained to fixed, small temporal context windows, and thus cannot consistently resolve aliasing ambiguities across the full duration of the video (illustrated in Fig. 2 in the case of synthesizing periodic motion, e.g., walking). (iii) Cascaded training regimens in general suffer from a domain gap, where the TSR model is trained on real downsampled video frames, but at inference time is used to interpolate generated frames, which accumulates errors.

Here, we take a different approach by introducing a new T2V diffusion framework that generates the full temporal duration of the video at once. We achieve this by using a Space-Time U-Net (STUNet) architecture that learns to downsample the signal in both space and time, and performs the majority of its computation in a compact space-time representation. This approach allows us to generate 80 frames at 16fps (or 5 seconds, which is longer than the average shot duration in most media (Cutting & Candan, 2015)) with a single base model, leading to more globally coherent motion compared to prior work. Surprisingly, this design choice has been overlooked by previous T2V models, which follow the convention to include only spatial down- and up-sampling operations in the architecture, and maintain a fixed temporal resolution across the network (Ho et al., 2022b, a; Singer et al., 2022; Ge et al., 2023; Blattmann et al., 2023b; Wang et al., 2023a; Guo et al., 2023; Zhang et al., 2023a; Girdhar et al., 2023; Po et al., 2023).

To benefit from the powerful generative prior of T2I models, we follow the trend of building Lumiere on top of a pretrained (and fixed) T2I model (Hong et al., 2022; Singer et al., 2022; Saharia et al., 2022b). In our case, the T2I model works in pixel space and consists of a base model followed by a spatial super-resolution (SSR) cascade. Since the SSR network operates at high spatial resolution, applying it on the entire video duration is infeasible in terms of memory requirements. Common SSR solutions use a temporal windowing approach, which splits the video into non-overlapping segments and stitches together the results. However, this can lead to inconsistencies in appearance at the boundaries between windows (Girdhar et al., 2023). We propose to extend Multidiffusion (Bar-Tal et al., 2023), an approach proposed for achieving global continuity in panoramic image generation, to the temporal domain, where we compute spatial super-resolution on temporal windows, and aggregate results into a globally coherent solution over the whole video clip.

We demonstrate state-of-the-art video generation results and show how to easily adapt Luimere to a plethora of video content creation tasks, including video inpainting (Fig. 7), image-to-video generation (Fig. 5), or generating stylized videos that comply with a given style image (Fig. 6). Finally, we demonstrate that generating the full video at once allows us to easily invoke off-the-shelf editing methods to perform consistent editing (Fig. 9).

Related work

Most of the common approaches for text-to-image (T2I) generation are based on diffusion models (Sohl-Dickstein et al., 2015; Ho et al., 2020; Song et al., 2020). Of these, DALL-E2 (Ramesh et al., 2022) and Imagen (Saharia et al., 2022b) achieve photorealistic text-to-image generation using cascaded diffusion models, whereas Stable Diffusion (Rombach et al., 2022) performs generation in a compressed low-dimensional latent space. A promising line of works design T2I diffusion models that generate high-resolution images end-to-end, without a spatial super-resolution cascaded system or fixed pre-trained latent space (Hoogeboom et al., 2023; Gu et al., 2023; Chen, 2023). Here, we design a T2V model that generates the full frame duration at once, avoiding the temporal cascade commonly involved in T2V models.

Text-to-Video Generation.

Recently, there have been substantial efforts in training large-scale T2V models on large scale datasets with autoregressive Transformers (e.g., (Villegas et al., 2023; Wu et al., 2022; Hong et al., 2022; Kondratyuk et al., 2023)) or Diffusion Models (e.g., (Ho et al., 2022a, b; Gupta et al., 2023)). A prominent approach for T2V generation is to “inflate” a pre-trained T2I model by inserting temporal layers to its architecture, and fine-tuning only those, or optionally the whole model, on video data (Singer et al., 2022; Blattmann et al., 2023b; Girdhar et al., 2023; Ge et al., 2023; Yuan et al., 2024). PYoCo (Ge et al., 2023) carefully design video noise prior and obtain better performance for fine-tuning a T2I model for video generation. VideoLDM (Blattmann et al., 2023b) and AnimateDiff (Guo et al., 2023) inflate StableDiffusion (Rombach et al., 2022) and train only the newly-added temporal layers, showing they can be combined with the weights of personalized T2I models. Interestingly, the ubiquitous convention of existing inflation schemes is to maintain a fixed temporal resolution across the network, which limits their ability to process full-length clips. In this work, we design a new inflation scheme which includes learning to downsample the video in both space and time, and performing the majority of computation in the compressed space-time feature space of the network. We extend an Imagen T2I model (Saharia et al., 2022b), however our architectural contributions could be used for latent diffusion as well, and are orthogonal to possible improvements to the diffusion noise scheduler (Ge et al., 2023) or to the video data curation (Blattmann et al., 2023a).

Lumiere

We utilize Diffusion Probabilistic Models as our generative approach (Sohl-Dickstein et al., 2015; Croitoru et al., 2023a; Dhariwal & Nichol, 2021; Ho et al., 2020; Nichol & Dhariwal, 2021). These models are trained to approximate a data distribution (in our case, a distribution over videos) through a series of denoising steps. Starting from a Gaussian i.i.d. noise sample, the diffusion model gradually denoises it until reaching a clean sample drawn from the approximated target distribution. Diffusion models can learn a conditional distribution by incorporating additional guiding signals, such as text embedding, or spatial conditioning (e.g., depth map) (Dhariwal & Nichol, 2021; Saharia et al., 2022a; Croitoru et al., 2023b; Zhang et al., 2023b).

Our framework consists of a base model and a spatial super-resolution (SSR) model. As illustrated in Fig. 3b, our base model generates full clips at a coarse spatial resolution. The output of our base model is spatially upsampled using a temporally-aware SSR model, resulting with the high-resolution video. We next describe the key design choices in our architecture, and demonstrate the applicability of our framework for a variety of downstream applications.

To make our problem computationally tractable, we propose to use a space-time U-Net which downsamples the input signal both spatially and temporally, and performs the majority of its computation on this compact space-time representation. We draw inspiration from Çiçek et al. (2016), who generalize the U-Net architecture (Ronneberger et al., 2015) to include 3D pooling operations for efficient processing of volumetric biomedical data.

Our architecture is illustrated in Fig. 4. We interleave temporal blocks in the T2I architecture, and insert temporal down- and up-sampling modules following each pre-trained spatial resizing module (Fig. 4a). The temporal blocks include temporal convolutions (Fig. 4b) and temporal attention (Fig. 4c). Specifically, in all levels except for the coarsest, we insert factorized space-time convolutions (Fig. 4b) which allow increasing the non-linearities in the network compared to full-3D convolutions while reducing the computational costs, and increasing the expressiveness compared to 1D convolutions (Tran et al., 2018). As the computational requirements of temporal attention scale quadratically with the number of frames, we incorporate temporal attention only at the coarsest resolution, which contains a space-time compressed representation of the video. Operating on the low dimensional feature map allows us to stack several temporal attention blocks with limited computational overhead.

Similarly to (Blattmann et al., 2023b; Guo et al., 2023), we train the newly added parameters, and keep the weights of the pre-trained T2I fixed. Notably, the common inflation approach ensures that at initialization, the T2V model is equivalent to the pre-trained T2I model, i.e., generates videos as a collection of independent image samples. However, in our case, it is impossible to satisfy this property due to the temporal down- and up-sampling modules. We empirically found that initializing these modules such that they perform nearest-neighbor down- and up- sampling operations results with a good starting point (see App. B).

2 Multidiffusion for Spatial-Super Resolution

The solution to this problem is given by linearly combining the predictions over overlapping windows. See App. C.

Applications

The lack of a TSR cascade makes it easier to extend Lumiere to downstream applications. In particular, our model provides an intuitive interface for downstream applications that require an off-the-shelf T2V model (e.g., Meng et al. (2022); Poole et al. (2023); Gal et al. (2023)). We demonstrate this property by performing video-to-video editing using SDEdit (Meng et al., 2022) (see Fig. 9). We next discuss a number of such applications, including style conditioned generation, image-to-video, inpainting and outpainting, and cinemagraphs. We present example frames in Figs. 6-9 and refer the reader to the Supplementary Material (SM) on our webpage for full video results.

Recall that we only train the newly-added temporal layers and keep the pre-trained T2I weights fixed. Previous work showed that substituting the T2I weights with a model customized for a specific style allows to generate videos with the desired style (Guo et al., 2023). We observe that this simple “plug-and-play” approach often results in distorted or static videos (see SM), and hypothesize that this is caused by the significant deviation in the distribution of the input to the temporal layers from the fine-tuned spatial layers.

Inspired by the success of GAN-based interpolation approaches (Pinkney & Adler, 2020), we opt to strike a balance between style and motion by linearly interpolating between the fine-tuned T2I weights, WstyleW_{\text{style}}, and the original T2I weights, WorigW_{\text{orig}}. Specifically, we construct the interpolated weights as Winterpolate=α⋅Wstyle+(1−α)⋅WorigW_{\text{interpolate}}=\alpha\cdot W_{\text{style}}+(1-\alpha)\cdot W_{\text{orig}}. The interpolation coefficient α∈[0.5,1]\alpha\in[0.5,1] is chosen manually in our experiments to generate videos that adhere to the style and depict plausible motion.

Figure 6 presents sample results for various styles from (Sohn et al., 2023). While more realistic styles such as “watercolor painting” result in realistic motion, other, less realistic spatial priors derived from vector art styles, result in corresponding unique non-realistic motion. For example, the “line drawing” style results in animations that resemble pencil strokes “drawing” the described scene, while the “cartoon” style results in content that gradually “pops out” and constructs the scene (see SM for full videos).

2 Conditional Generation

In this case, the first frame of the video is given as input. The conditioning signal CC contains this first frame followed by blank frames for the rest of the video. The corresponding mask MM contains ones (i.e., unmasked content) for the first frame and zeros (i.e., masked content) for the rest of the video. Figures 1 and 5 show sample results of image-conditioned generation (see SM for more results). Our model generates videos that start with the desired first frame, and exhibit intricate coherent motion across the entire video duration.

Inpainting.

Here, the conditioning signals are a user-provided video CC and a mask MM that describes the region to complete in the video. Note that the inpainting application can be used for object replacement/insertion (Fig. 1) as well as for localized editing (Fig. 7). The effect is a seamless and natural completion of the masked region, with contents guided by the text prompt. We refer the reader to the SM for more examples of both inpainting and outpainting.

Cinemagraphs.

We additionally consider the application of animating the content of an image only within a specific user-provided region. The conditioning signal CC is the input image duplicated across the entire video, while the mask MM contains ones for the entire first frame (i.e., the first frame is unmasked), and for the other frames, the mask contains ones only outside the user-provided region (i.e., the other frames are masked inside the region we wish to animate). We provide sample results in Fig. 8 and in the SM. Since the first frame remains unmasked, the animated content is encouraged to maintain the appearance from the conditioning image.

Evaluation and Comparisons

We train our T2V model on a dataset containing 30M videos along with their text caption. The videos are 80 frames long at 16 fps (5 seconds). The base model is trained at 128×128128\times 128 and the SSR outputs 1024×10241024\times 1024 frames. We evaluate our model on a collection of 109 text prompts describing diverse objects and scenes. The prompt list consists of 91 prompts used by prior works (Singer et al., 2022; Ho et al., 2022a; Blattmann et al., 2023b) and the rest were created by us (see App. D). Additionally, we employ a zero-shot evaluation protocol on the UCF101 dataset (Soomro et al., 2012), as detailed in Sec. 5.2.

We illustrate text-to-video generation in Figs. 1 and 5. Our method generates high-quality videos depicting both intricate object motion (e.g., walking astronaut in Fig. 5) and coherent camera motion (e.g., car example in Fig. 1). We refer the reader to the SM for full-video results.

We compare our method to prominent T2V diffusion models: (i) ImagenVideo (Ho et al., 2022a), that operates in pixel-space and consists of a cascade of 7 models (a base model, 3 TSR models, and 3 SSR models); (ii) AnimateDiff (Guo et al., 2023), (iii) StableVideoDiffusion (SVD) (Blattmann et al., 2023a), and (iv) ZeroScope (Wang et al., 2023a) that inflate Stable Diffusion (Rombach et al., 2022) and train on video data; note that AnimateDiff and ZeroScope output only 16, and 36 frames respectively. SVD released only their image-to-video model, which outputs 25 frames and is not conditioned on text. Additionally, we compare to (v) Pika (Pika labs, 2023) and (vi) Gen-2 (RunwayML, 2023) commercial T2V models that have available API. Furthermore, we quantitatively compare to additional T2V models that are closed-source in Sec. 5.2.

1 Qualitative Evaluation

We provide qualitative comparison between our model and the baselines in Fig. 11. We observed that Gen-2 (RunwayML, 2023) and Pika (Pika labs, 2023) demonstrate high per-frame visual quality; however, their outputs are characterized by a very limited amount of motion, often resulting in near-static videos. ImagenVideo (Ho et al., 2022a) produces a reasonable amount of motion, but at a lower overall visual quality. AnimateDiff (Guo et al., 2023) and ZeroScope (Wang et al., 2023a) exhibit noticeable motion but are also prone to visual artifacts. Moreover, they generate videos of shorter durations, specifically 2 seconds and 3.6 seconds, respectively. In contrast, our method produces 5-second videos that have higher motion magnitude while maintaining temporal consistency and overall quality.

2 Quantitative Evaluation

Following the evaluation protocols of Blattmann et al. (2023a) and Ge et al. (2023), we quantitatively evaluate our method for zero-shot text-to-video generation on UCF101 (Soomro et al., 2012). Table 1 reports the Fréchet Video Distance (FVD) (Unterthiner et al., 2018) and Inception Score (IS) (Salimans et al., 2016) of our method and previous work. We achieve competitive FVD and IS scores. However, as discussed in previous work (e.g., Girdhar et al. (2023); Ho et al. (2022a); Chong & Forsyth (2020)), these metrics do not faithfully reflect human perception, and may be significantly influenced by low-level details (Parmar et al., 2022) and by the distribution shift between the reference UCF101 data and the T2V training data (Girdhar et al., 2023). Furthermore, the protocol uses only 16 frames from generated videos and thus is not able to capture long-term motion.

User Study.

We adopt the Two-alternative Forced Choice (2AFC) protocol, as used in previous works (Kolkin et al., 2019; Zhang et al., 2018; Blattmann et al., 2023a; Rombach et al., 2022). In this protocol, participants were presented with a randomly selected pair of videos: one generated by our model and the other by one of the baseline methods. Participants were then asked to choose the video they deemed better in terms of visual quality and motion. Additionally, they were asked to select the video that more accurately matched the target text prompt. We collected ∼\sim400 user judgments for each baseline and question, utilizing the Amazon Mechanical Turk (AMT) platform. As illustrated in Fig. 10, our method was preferred over all baselines by the users and demonstrated better alignment with the text prompts. Note that ZeroScope and AnimateDiff generate videos only at 3.6 and 2 second respectively, we thus trim our videos to match their duration when comparing to them.

We further conduct a user study for comparing our image-to-video model (see Sec. 4.2) against Pika (Pika labs, 2023), StableVideoDiffusion (SVD) (Blattmann et al., 2023a), and Gen2(RunwayML, 2023). Note that SVD image-to-video model is not conditioned on text, we thus focus our survey on the video quality. As seen in Fig. 10, our method was preferred by users compared to the baselines. For a detailed description of the full evaluation protocol, please refer to Appendix D.

Conclusion

We presented a new text-to-video generation framework, utilizing a pre-trained text-to-image diffusion model. We identified an inherent limitation in learning globally-coherent motion in the prevalent approach of first generating distant keyframes and subsequently interpolating them using a cascade of temporal super-resolution models. To tackle this challenge, we introduced a space-time U-Net architecture design that directly generates full-frame-rate video clips, by incorporating both spatial, and temporal down- and up-sampling modules. We demonstrated state-of-the-art generation results, and showed the applicability of our approach for a wide range of applications, including image-to-video, video inapainting, and stylized generation.

As for limitations, our method is not designed to generate videos that consist of multiple shots, or that involve transitions between scenes. Generating such content remains an open challenge for future research. Furthermore, we established our model on top of a T2I model that operates in the pixel space, and thus involves a spatial super resolution module to produce high resolution images. Nevertheless, our design principles are applicable to latent video diffusion models (Rombach et al., 2022), and can trigger further research in the design of text-to-video models.

Societal Impact

Our primary goal in this work is to enable novice users to generate visual content in a creative and flexible way. However, there is a risk of misuse for creating fake or harmful content with our technology, and we believe that it is crucial to develop and apply tools for detecting biases and malicious use cases in order to ensure a safe and fair use.

We would like to thank Ronny Votel, Orly Liba, Hamid Mohammadi, April Lehman, Bryan Seybold, David Ross, Dan Goldman, Hartwig Adam, Xuhui Jia, Xiuye Gu, Mehek Sharma, Rachel Hornung, Oran Lang, Jess Gallegos, William T. Freeman and David Salesin for their collaboration, helpful discussions, feedback and support. We thank owners of images and videos used in our experiments for sharing their valuable assets (attributions can be found in our webpage).

References

Appendix

Appendix A Qualitative Comparison

We provide a qualitative comparison of our method and the baselines (see Sec. 5.1).

Appendix B Initialization

As illustrated in Fig. 4(b)-(c), each inflation block has a residual structure, in which the temporal components are followed by a linear projection, and are added to the output of the pre-trained T2I spatial layer. Similarly to previous works (e.g., (Guo et al., 2023)), we initialize the linear projection to a zero mapping, such that at initialization, the inflation block is equivalent to the pre-trained T2I block. However, as mentioned in Sec. 3.1, due to the temporal down- and up- sampling modules, the overall inflated network cannot be equivalent at initialization to the pre-trained T2I model due to the temporal compression. We empirically found that initializing the temporal down- and up-sampling modules such that they perform nearest-neighbor down- and up- sampling operations (i.e., temporal striding or frame duplication) is better than a standard initialization (He et al., 2015). In more detail, the temporal downsampling is implemented as a 1D temporal convolution (identity initialized) followed by a stride, and the temporal upsampling is implemented as frame duplication followed by a 1D temporal convolution (identity initialized). Note that this scheme ensures that at initialization, every NthN^{\text{th}} frame is identical to the output of the pre-trained T2I model, where NN is the total temporal downsampling factor in the network (see Fig. 13). We ablate our choice of temporal down- and up- sampling initialization on the UCF-101 (Soomro et al., 2012) dataset in Fig. 12.

Appendix C SSR with MultiDiffusion

As discussed in Sec. 3.2, we apply MultiDiffusion (Bar-Tal et al., 2023) along the temporal axis in order to feed segments of the video to the SSR network. Specifically, the temporal windows have an overlap of 2 frames, and at each denoising step we average the predictions of overlapping pixels. MultiDiffusion allows us to avoid temporal boundary artifacts between segments of the video (see Fig. 14).

Appendix D Evaluation Protocol

We used a set of 109 text prompts containing a variety of objects and actions for the text-to-video and image-to-video surveys. 9191 prompts were taken from various recent text-to-video generation methods (Singer et al., 2022; Ho et al., 2022a; Blattmann et al., 2023b). The rest are prompts we created which describe complex scenes and actions, see Sec. D.3 for the full list of prompts. For each baseline method, we generated results for all prompts through their official APIs. To ensure a fair comparison, when generating results using our method we’ve fixed a random seed and generated all results using the same random seed without any curation. In addition, following (Girdhar et al., 2023) we align the spatial resolution and temporal duration of our videos to each one of the baselines. Spatially, we apply a central crop followed by resizing to 512×\times512, and temporally, by trimming the start and the end such that the duration of both videos matches the shortest one. Each participant is shown 10 side-by-side comparisons between our method and a baseline (randomly ordered). We then asked the participant ”Which video matches the following text?” and ”Which video is better? Has more motion and better quality”. We enable the user to answer only after watching the full video, waiting at least 3 seconds, and pressing a button reading “Ready to answer” (Fig. 15). To ensure credible responses, we include several vigilance tests during the study - a real video vs a complete random noise, and a static image vs a real video.

D.2 Zero-shot evaluation on UCF101

To compute Fréchet Video Distance (FVD), we generated 10,235 videos by following the class distribution of the UCF101 dataset. We use the same prompt set from Ge et al. (2023) to generate the videos. After resizing all videos to 244×\times244 resolution, we extract I3D embedding (Carreira & Zisserman, 2017) of the first 16 frames of our videos. Then we compute the FVD score between the I3D embedding of ours and that of UCF101 videos. To compute Inception Score (IS), the same generated videos are used to extract C3D embedding (Saito et al., 2020).

D.3 List of prompts used for the user study evaluation

”A bear is giving a presentation in classroom.”

”A bear wearing sunglasses and hosting a talk show.”

”A beautiful sunrise on mars, Curiosity rover. High definition, timelapse, dramatic colors.”

”A beautiful sunrise on mars. High definition, timelapse, dramatic colors.”

”A big moon rises on top of Toronto city.”

”A car moving slowly on an empty street, rainy evening, van Gogh painting.”

”A cat wearing sunglasses and working as a lifeguard at a pool.”

”A confused grizzly bear in calculus class.”

”A cute happy Corgi playing in park, sunset, 4k.”

”A dog driving a car on a suburban street wearing funny sunglasses”

”A dog wearing a Superhero outfit with red cape flying through the sky.”

”A dog wearing virtual reality goggles in sunset, 4k, high resolution.”

”A fantasy landscape, trending on artstation, 4k, high resolution.”

”A fat rabbit wearing a purple robe walking through a fantasy landscape.”

”A fire dragon breathing, trending on artstation, slow motion.”

”A glass bead falling into water with a huge splash. Sunset in the background.”

”A golden retriever has a picnic on a beautiful tropical beach at sunset, high resolution.”

”A goldendoodle playing in a park by a lake.”

”A happy elephant wearing a birthday hat walking under the sea.”

”A horse galloping through van Gogh’s Starry Night.”

”A hot air balloon ascending over a lush valley, 4k, high definition.”

”A koala bear playing piano in the forest.”

”A lion standing on a surfboard in the ocean in sunset, 4k, high resolution.”

”A monkey is playing bass guitar, stage background, 4k, high resolution.”

”A panda standing on a surfboard in the ocean in sunset, 4k, high resolution.”

”A polar bear is playing bass guitar in snow, 4k, high resolution.”

”A raccoon dressed in suit playing the trumpet, stage background, 4k, high resolution.”

”Red sports car coming around a bend in a mountain road.”

”A shark swimming in clear Carribean ocean.”

”A shiny golden waterfall flowing through glacier at night.”

”A shooting star flying through the night sky over mountains.”

”A sloth playing video games and beating all the high scores.”

”A small hand-crafted wooden boat taking off to space.”

”A stunning aerial drone footage time lapse of El Capitan in Yosemite National Park at sunset.”

”A swarm of bees flying around their hive.”

”A teddy bear is playing the electric guitar, high definition, 4k.”

”A video of the Earth rotating in space.”

”A video that showcases the beauty of nature, from mountains and waterfalls to forests and oceans.”

”Aerial view of a hiker man standing on a mountain peak.”

”Aerial view of a snow-covered mountain.”

”Aerial view of a white sandy beach on the shores of a beautiful sea, 4k, high resolution.”

”An animated painting of fluffy white clouds moving in sky.”

”An astronaut cooking with a pan and fire in the kitchen, high definition, 4k.”

”An astronaut riding a horse, high definition, 4k.”

”An astronaut riding a horse in sunset, 4k, high resolution.”

”Artistic silhouette of a wolf against a twilight sky, embodying the spirit of the wilderness..”

”Back view on young woman dressed in a yellow jacket walking in the forest.”

”Maltipoo dog on the carpet next to a Christmas tree in a beautiful living room.”

”Beer pouring into glass, low angle video shot.”

”Bird-eye view of a highway in Los Angeles.”

”Campfire at night in a snowy forest with starry sky in the background.”

”Cherry blossoms swing in front of ocean view, 4k, high resolution.”

”Chocolate syrup pouring on a vanilla ice cream”

”Close up of grapes on a rotating table. High definition.”

”Teddy bear surfer rides the wave in the tropics.”

”Drone flythrough interior of sagrada familia cathedral.”

”Drone flythrough of a fast food restaurant on a dystopian alien planet.”

”Drone flythrough of a tropical jungle covered in snow.”

”Edge lit hamster writing code on a computer in front of a scifi steampunk machine with yellow green red and blue gears high quality dslr professional photograph”

”Epic tracking shot of a Powerful silverback gorilla male walking gracefully”,

”Flying through a temple in ruins, epic, forest overgrown, columns cinematic, detailed, atmospheric, epic, mist, photo-realistic, concept art, volumetric light, cinematic epic, 8k”

”Flying through an intense battle between pirate ships in a stormy ocean.”

”Flying through fantasy landscapes, 4k, high resolution.”

”Pug dog listening to music with big headphones.”

”Horror house living room interior overview design, 8K, ultra wide angle, pincushion lens effect.”

”Incredibly detailed science fiction scene set on an alien planet, view of a marketplace. Pixel art.”

”Jack russell terrier dog snowboarding. GoPro shot.”

”Low angle of pouring beer into a glass cup.”

”Melting pistachio ice cream dripping down the cone.”

”Milk dripping into a cup of coffee, high definition, 4k.”

”Pouring latte art into a silver cup with a golden spoon next to it.

”Studio shot of minimal kinetic sculpture made from thin wire shaped like a bird on white background.”

”Sunset time lapse at the beach with moving clouds and colors in the sky, 4k, high resolution.”

”Teddy bear walking down 5th Avenue, front view, beautiful sunset, close up, high definition, 4k.”

”The Orient Express driving through a fantasy landscape, animated oil on canvas.”

”Time lapse at a fantasy landscape, 4k, high resolution.”

”Time lapse at the snow land with aurora in the sky, 4k, high resolution.”

”Tiny plant sprout coming out of the ground.”

”Toy Poodle dog rides a penny board outdoors”

”Traveler walking alone in the misty forest at sunset.”

”Two pandas discussing an academic paper.”

”Two pandas sitting at a table playing cards, 4k, high resolution.”

”Two raccoons are playing drum kit in NYC Times Square, 4k, high resolution.”

”Two raccoons reading books in NYC Times Square.”

”US flag waving on massive sunrise clouds in bad weather. High quality, dynamic, detailed, cinematic lighting”

”View of a castle with fantastically high towers reaching into the clouds in a hilly forest at dawn.”

”Waves crashing against a lone lighthouse, ominous lighting.”

”Woman in white dress waving on top of a mountain.”

”Wooden figurine surfing on a surfboard in space.”