WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion Model
Zongjian Li, Bin Lin, Yang Ye, Liuhan Chen, Xinhua Cheng, Shenghai Yuan, Li Yuan
Introduction
The release of Sora , a video generation model developed by OpenAI, has pushed the boundaries of synthesizing photorealistic videos, drawing unprecedented attention to the field of video generation. Recent advancements in Latent Video Diffusion Models (LVDMs), such as Open-Sora Plan , Open-Sora , CogVideoX , EasyAnimate , Movie Gen , and Allegro , have led to substantial improvements in video generation quality. These methods establish a compressed latent space using a pretrained video Variational Autoencoder (VAE), where the generative performance is fundamentally determined by the compression quality.
Current video VAEs remain constrained by fully convolutional architectures inherited from image era. They attempt to address video flickering and redundant information by incorporating spatio-temporal interaction layers and spatio-temporal compression layers. Several recent works, including OD-VAE , CogVideoX , CV-VAE , and Allegro adopt dense 3D structure to achieve high-quality video compression. While these methods demonstrate impressive reconstruction performance, they require prohibitively intensive computational resources. In contrast, alternative approaches such as Movie Gen , and Open-Sora utilize 2+1D architecture, resulting in reduced computational requirements at the cost of lower reconstruction quality. Previous architectures have not fully exploited video temporal redundancy, leading to requirements for redundant spatio-temporal interaction layers to enhance video compression quality.
Many attempts employ block-wise inference strategies to trade computation time for memory, thus addressing computational bottlenecks in processing high-resolution, long-duration videos. EasyAnimate introduced Slice VAE to encode and decode video frames in groups, but this approach leads to discontinuous output videos. Several methods, including Open-Sora , Open-Sora Plan , Allegro , and Movie Gen , implement tiling inference strategies, but often produce spatio-temporal artifacts in overlapping regions. Although CogVideoX employ caching to ensure convolution continuity, its reliance on group normalization disrupts the independence of temporal feature, thus preventing lossless block-wise inference. These limitations highlight two core challenges faced by video VAEs: (1) excessive computational demands due to redundant architecture, and (2) compromised latent space integrity resulting from existing tiling inference strategies, which cause artifacts and flickering in reconstructed videos.
Wavelet transform decomposes videos into multiple frequency-domain components. This decomposition enables prioritization strategies for encoding crucial video components. In this work, we propose Wavelet Flow VAE (WF-VAE), a novel autoencoder that utilizes multi-level wavelet transforms for extracting multi-scale pyramidal features and establishes a main energy flow pathway for these features to flow into latent representation. This pathway bypasses low-frequency video information to latent space, skipping the backbone network. Our WF-VAE enables a simplified backbone design with reduced 3D convolutions, significantly reducing computational costs. To address the potential latent space disruption, we propose Causal Cache mechanism. This approach leverages the properties of causal convolution and maintains the continuity of the convolution sliding window through a caching strategy, which ensuring numerical identity between block-wise inference and direct inference results. Experimental results show that WF-VAE achieves state-of-the-art performance in both reconstruction quality and computational efficiency. To summarize, major contributions of our work include:
We propose WF-VAE, which leverages multi-level wavelet transforms to extract pyramidal features and establishes a main energy flow pathway for video information flow into a latent representation.
We introduce a lossless block-wise inference mechanism called Causal Cache which maintains identical performance as direct inference across videos of any duration.
Extensive experimental evaluations of video reconstruction and generation demonstrate that WF-VAE achieves state-of-the-art performance, as validated through comprehensive ablation studies.
Related Work
Variational Autoencoders. introduced the VAE based on variational inference, establishing a novel generative network structure. Subsequent research demonstrated that training and inference in VAE latent space could substantially reduce computational costs for diffusion models. further proposed a two-stage image synthesis approach by decoupling perceptual compression from diffusion model. After that, numerous studies explored video VAEs with a focus on more efficient video compression including Open-Sora Plan , CogVideoX , and other models . Current video VAE architectures largely derive from earlier image VAE design and use a convolutional backbone. To address the challenges of inference on large videos, current video VAEs employ block-wise inference.
Latent Video Diffusion Models. In the early stages of Latent Video Diffusion Models (LVDMs) development, models like AnimateDiff , and MagicTime primarily utilized the U-Net backbone for denoising, without temporal compression in VAEs. Following the paradigm introduced by Sora, recent open-source models such as Open-Sora , Open-Sora Plan , and CogVideoX have adopted DiT backbone with a spatiotemporally compressed VAE. Some methods including Open-Sora and Movie Gen , employ a 2.5D design in either the DiT backbone or the VAE to reduce training costs. For LVDMs, the upper limit of video generation quality is largely determined by the VAE’s reconstruction quality. As overhead increase rapidly in scale, optimizing the VAE becomes crucial for handling large-scale data and enabling extensive pre-training.
Method
Preliminary. The Haar wavelet transform , a fundamental form of wavelet transform, is widely used in signal processing . It efficiently captures spatio-temporal information by decomposing signals through two complementary filters. The first is the Haar scaling filter , which acts as a low-pass filter that captures the average or approximation coefficients. The second is the Haar wavelet filter , which functions as a high-pass filter that extracts the detail coefficients. These orthogonal filters are designed to be simple yet effective, with the scaling filter smoothing the signal and the wavelet filter detecting local changes or discontinuities.
where represent the filters applied along each dimension, and represents the convolution operation. The transform begins with , and for subsequent layers, , indicating that each layer operates on the low-frequency component from the previous layer. At each decomposition layer , the transform produces eight sub-band components: . Here, represents the low-frequency component across all dimensions, while captures high-frequency details. To implement different downsampling rates in the temporal and spatial dimensions, a combination of 2D and 3D wavelet transforms can be implemented. Specifically, to obtain a compression rate of 4×8×8 (temporal×height×width), we can employ a combination of two-layer 3D wavelet transform followed by one-layer 2D wavelet transform.
2 Architecture Design of WF-VAE
Through analyzing different sub-bands, we found that video energy is mainly concentrated in the low-frequency sub-band . Based on this observation, we establish an energy flow pathway outside the backbone, as show in Fig. 2, so that low-frequency information can smoothly flow from video to latent representation during encoding process, and then flow back to the video during the decoding process. This would inherently allow the model to attention more on low-frequency information, and apply higher compression rates to high-frequency information. With the additional path, we can reduce the computational cost brought by the dense 3D convolutions in the backbone.
Specifically, given a video , we apply multi-level wavelet transform to obtain pyramid features and . We utilize as the input for the encoder and the target output for the decoder. The convolutional backbone and multi-level wavelet transform layer downsample the feature maps simultaneously at every downsampling layer, enabling the concatenation of features from two branches in the backbone. We employ Inflow Block to transform the channel numbers of and to , which are then concatenated with feature maps from backbone. We compare to the width of the energy flow pathway, as analyzed in Sec. 4.4. On the decoder side, we maintain a structure symmetrical to the encoder. We extract feature maps with channels from the backbone and process them through Outflow Block to obtain . In order to allow information to flow from the lower layer to the sub-band of the next layer, we have:
Overall, we create shortcuts for low-frequency information, ensuring its greater emphasis in latent representation.
3 Causal Cache
We replace regular 3D convolutions with causal 3D convolutions in WF-VAE. The causal convolution applies temporal padding at the start with kernel size . This padding strategy ensures the first frame remains independent from subsequent frames, thus enabling the processing of images and videos within a unified architecture. Furthermore, we leverage the causal convolution properties to achieve lossless inference. We first extract the initial frame from a video with frames. The remaining frames are then partitioned into temporal chunks, where represents the chunk size. Let denote the temporal convolutional stride, and represents the chunk block index. To maintain the continuity of convolution sliding windows, each chunk caches its tail frames for the next chunk. The number of cached frames is given by:
For example, , the equation yields , as illustrated in Fig. 9 (a). Similarly, with , we obtain , indicating only the last frame requires to be cached. Special cases exist, such as when , which results in . Fig. 9(b) provides a qualitative comparison between Causal Cache and the tiling strategy, illustrating how it effectively mitigates significant distortions in both color and shape.
4 Training Objective
Following the training strategies of , our loss function combines multiple components, including reconstruction loss (comprising L1 and perceptual loss ), adversarial loss, and KL regularization . Our model is characterized by a low-frequency energy flow and symmetry between the encoder and decoder. To maintain this architectural principle, we introduce a regularization term denoted as (WL loss), which enforces structural consistency by penalizing deviations from the intended energy flow:
The final loss function is formulated as:
We examine the impact of the weighting factor in Sec. 4.4. Following , we implement dynamic adversarial loss weighting to balance the relative gradient magnitudes between adversarial and reconstruction losses:
where denotes the gradient with respect to last layer of decoder, and is used for numerical stability.
Experiments
Baseline Models. To assess the effectiveness of WF-VAE, we perform a comprehensive evaluation, comparing its performance and efficiency against several state-of-the-art VAE models. The models considered are: (1) OD-VAE , a 3D causal convolutional VAE used in Open-Sora Plan 1.2 ; (2) Open-Sora VAE ; (3) CV-VAE ; (4) CogVideoX VAE ; (5) Allegro VAE ; (6) SVD-VAE , which does not compress temporally and (7) SD-VAE , a widely used image VAE. Among these, CogVideoX VAE has a latent dimension of 16, while others use a latent dimension of 4. Notably, all VAEs except CV-VAE and SD-VAE, have been validated on LVDMs previously, making them highly representative for comparison.
Dataset & Evaluation. We utilize the Kinetics-400 dataset for both training and validation. For testing, we employ the Panda70M and WebVid-10M datasets. To comprehensively evaluate the model’s reconstruction performance, we select Peak Signal-to-Noise Ratio (PSNR) , Learned Perceptual Image Patch Similarity (LPIPS) , and Structural Similarity Index Measure (SSIM) as primary evaluation metrics. Additionally, we use Fréchet Video Distance (FVD) to assess visual quality and temporal coherence. To assess our model’s performance in generating results with the diffusion model, we utilize the UCF-101 and SkyTimelapse datasets for conditional and unconditional training 100,000 steps. Following , we extract 16-frame clips of 2,048 videos to compute FVD16. Additionally, we evaluate the Inception Score (IS) exclusively on the UCF-101 dataset, as suggested by . We select Latte-L as the denoiser. Since our primary focus is not on the final generative performance but rather on whether the latent spaces of various video VAEs facilitate effective diffusion model training, we chose not to use the higher-performing Latte-XL.
Training Strategy. We employ the AdamW optimizer with parameters and , and set a fixed learning rate of . Our training process comprises three stages: (I) the first stage aligns with , where we preprocess videos to 25 frames with resolution, and a total batch size of 8. (II) we refresh the discriminator, increase the number of frames to 49, and reduce the FPS by half to enhance motion dynamics. (III) we observe that a large significantly affects video stability; therefore, we refresh the discriminator once more and set to . All three stages employ L1 loss: the initial stage is trained for 800,000 steps, while the subsequent stages are each trained for 200,000 steps. The training process utilizes 8 NVIDIA H100 GPUs. We implement a 3D discriminator and initiate GAN training from the start. All training hyperparameters are detailed in the appendix.
2 Comparison With Baseline Models
We compare WF-VAE with baseline models in three key aspects: computational efficiency, reconstruction performance, and performance in diffusion-based generation. To ensure fairness in comparing metrics and model efficiency, we disable block-wise inference strategies across all VAEs.
Computational Efficiency. The computational efficiency evaluations are conducted using an H100 GPU with float32 precision. Performance evaluations are performed at 33 frames across multiple input resolutions. Allegro VAE is evaluated using 32-frame videos to maintain consistency due to its non-causal convolution architecture. For a fair comparison, all benchmark VAEs do not employ block-wise inference strategies and are directly inferred. As shown in Fig. 4, WF-VAE demonstrates superior inference performance compared to other VAEs. For instance, WF-VAE-L requires 7170 MB of memory for encoding a video at 512×512 resolution, whereas OD-VAE demands approximately 31944 MB (445% higher). Similarly, CogVideoX consumes around 35849.33 MB (499% higher), and Allegro VAE requires 55664 MB (776% higher). These results highlight WF-VAE’s significant advantages in large-scale training and data processing. In terms of encoding speed, WF-VAE-L achieves an encoding time of 0.0513 seconds, while OD-VAE, CogVideoX, and Allegro VAE exhibit encoding times of 0.0945 seconds, 0.1810 seconds, and 0.3731 seconds, respectively, which are approximately 184%, 352%, and 727% slower. From Fig. 4, it can also be observed that WF-VAE has significant advantages in decoding computational efficiency.
Video Reconstruction Performance. We present a quantitative comparison of reconstruction performance between WF-VAE and baseline models in Tab. 1 and qualitative reconstruction results in Fig. 5. As shown, despite having the lowest computational cost, WF-VAE-S outperforms popular open-source video VAEs such as OD-VAE and Open-Sora VAE on both datasets. When increasing the model complexity, WF-VAE-L competes well with Allegro , outperforming it in PSNR, LPIPS and FVD, but slightly lagging in SSIM. However, WF-VAE-L is significantly more computationally efficient than Allegro. Additionally, we compare WF-VAE-L with CogVideoX using 16 latent channels. Except for a slightly lower SSIM on the Webvid-10M, WF-VAE-L outperform CogVideoX across all other metrics. These results indicate that WF-VAE achieves competitive reconstruction performance compared to state-of-the-art open-source video VAEs, while substantially improving computational efficiency.
Video Generation Evaluation. Fig. 6 and Tab. 2 qualitatively and quantitatively demonstrate the video generation results of the diffusion model using WF-VAE. WF-VAE achieves the best performance in terms of FVD and IS metrics. For instance, in the SkyTimelapse dataset, among models with 4 latent channels, WF-VAE-S achieves the best FVD score, which is 10.23 lower than WF-VAE-L and 27.35 lower than OD-VAE. For models with 16 latent channels, WF-VAE-L’s FVD score is 0.51 lower than CogVideoX’s. This might be because the higher dimensionality of the latent space makes convergence more difficult, resulting in slightly inferior performance compared to WF-VAE with 4 latent channels under the same training steps.
3 Ablation Study
Increasing the latent dimension. Recent works show that the number of latent channel significantly impacts reconstruction quality. We conduct experiments with 4, 8, 16, and 32 latent channels. Fig. 7(a) illustrates that reconstruction performance improves significantly as the number of latent channels increases. However, larger latent channels may increase convergence difficulty in training diffusion model, as evidenced by the results presented in Tab. 2.
Exploration of WL loss weight . To ensure structural symmetry in WF-VAE, we introduce WL loss. As shown in Fig. 7(b), the model exhibits a substantial decrease in PSNR performance when . Our experiments demonstrate optimal results for both PSNR and LPIPS metrics when . Furthermore, our analysis of different loss functions reveals that utilizing L1 loss produces superior results compared to L2 loss.
Expanding the energy flow path. The low-frequency information flows into the latent representation through the energy flow pathway, and the number of channels in this path, , determines the intensity of low-frequency information injection into the backbone. We conducted experiments with values of 64, 128, and 256. As shown in Fig. 7(c), we find that when is 128, it can balance reconstruction and computational performance.
Increasing the number of base channels. To further exploit the capabilities of the WF-VAE architecture, we increase the model complexity by expanding the number of base channels. The channel dimensionality increases by one base channel width after each downsampling layer, starting from the number of base channels. As shown in Tab. 3, we experiment with three configurations: 128, 160, and 192 base channels. The results demonstrate that the model’s performance improves as the number of base channels increases. Despite the corresponding rise in model parameters, the computational cost remains comparatively low compared to benchmark models, as shown in Fig. 4.
Ablation studies of model architecture. First, We evaluate the effectiveness of the energy pathway by examining its impact when removed from layer 3 alone and both layers 2 and 3. This analysis demonstrates the benefits of incorporating low-frequency video energy into the backbone. Second, we investigate the significance of our proposed WL Loss in regularizing the encoder-decoder Third, we examine the impact of modifying the normalization method to layer normalization for Causal Cache. The results of these ablation studies are shown in Tab. 4.
4 Causal Cache
To validate the lossless inference capability of Causal Cache, we compared our approach with existing block-wise inference methods implemented in several open-source LVDMs. OD-VAE and Allegro provide spatio-temporal tiling inference implementations, and CogVideoX implements a temporal caching strategy. As demonstrated in Tab. 5, both tiling strategies and conventional caching methods exhibited performance degradation, while Causal Cache achieves lossless inference with performance metrics identical to direct inference.
Conclusion
In this paper, we propose WF-VAE, an innovative autoencoder that utilizes multi-level wavelet transform to extract pyramidal features, thus creating a primary energy flow pathway for encoding low-frequency video information into a latent representation. Additionally, we introduce a lossless block-wise inference mechanism called Causal Cache, which completely resolves video flickering associated with prior tiling strategies. Our experiments demonstrate that WF-VAE achieves state-of-the-art reconstruction performance while maintaining low computational costs. WF-VAE significantly reduces the expenses associated with large-scale video pre-training, potentially inspiring future designs of video VAEs.
Limitations and Future Work. The initial design of the decoder incorporated insights from , employing a highly complex structure that resulted in more parameters in the backbone of the decoder compared to the encoder. Although the computational cost remains manageable, we consider these parameters redundant. Consequently, we aim to streamline the model in future work to fully leverage the advantages of our architecture.
References
Notations
The notations and their descriptions in the paper are shown in Tab. 6.
Wavelet Subband Analysis
We analyze the energy and entropy distributions across the subbands obtained after wavelet transform. As illustrated in Fig. 8(b), the energy and entropy of the video are primarily concentrated in the low-frequency subband. This concentration suggests that low-frequency components carry more significant information and necessitate lower compression rates to ensure superior reconstruction performance. This observation further validates the rationale behind our proposed approach.
Training Details
The training hyperaparameters are shown in Tab. 7.
Derivation of Causal Cache
where denotes the chunk size. For a given chunk index , the maximum sliding index is determined by the constraint :
Consequently, the required cache size for chunk is: