FlowTok: Flowing Seamlessly Across Text and Image Tokens

Ju He, Qihang Yu, Qihao Liu, Liang-Chieh Chen

Introduction

Bridging different modalities is essential for comprehending the diverse forms of data that represent our world, encompassing both understanding and generation. In multimodal understanding, extensive research has focused on designing architectures that project different modalities into a shared latent space . These approaches have significantly advanced cross-modal representation learning and real-world understanding by leveraging a common latent space between modalities.

In contrast, multimodal generation (e.g., text-to-image generation) follows a different paradigm, primarily relying on the diffusion process , where the source modality (e.g., text) serves as a conditioning signal to guide the denoising process. Various conditioning mechanisms have been explored, including concatenation , cross-attention , conditioning embeddings , and hybrid strategies . While effective, these approaches introduce substantial complexity, requiring intricate conditioning mechanisms and noise scheduling. This naturally raises an important question: Can we unify multimodal understanding and generation by enabling direct transitions within a shared latent space?

To address this question, we revisit flow matching —a modern generative framework that learns a direct path from noise to data, enabling faster convergence and accelerated sampling, leading to state-of-the-art multimodal generation results . Unlike diffusion models, flow matching is not constrained to using noise as the source distribution; instead, it only requires the source and target distributions to share the same shape. Pioneering works have demonstrated its effectiveness in learning direct mappings within the same modality (e.g., image-to-image generation). Meanwhile, CrossFlow extends flow matching to cross-modal learning by mapping text into a 2D latent space to match the shape of image embeddings, paving the way for new possibilities. However, while this approach simplifies the overall pipeline, it still operates on 2D latent representations. As a result, the additional computational overhead introduced by the text variational autoencoder in CrossFlow makes it slower than modern text-to-image diffusion models like SD1.5 and SD2.1 , ultimately contradicting its original goal of efficiency.

To this end, we introduce FlowTok, a minimal framework that enables seamless Flowing of Tokens across text and image—the two most prevalent modalities (Fig.˜2). At the core of FlowTok, both text and images are encoded into compact 1D latent tokens within a unified space, enabling direct flow matching between them. On the text side, FlowTok employs a pre-trained text encoder to extract initial 1D text embeddings. Since these embeddings typically reside in a higher-dimensional space than image latents, FlowTok introduces a lightweight text projector to map text embeddings into a low-dimensional variational latent space. On the image side, FlowTok builds on recent advancements in image tokenization to encode images into compact 1D latent tokens. Specifically, we enhance TA-TiTok by integrating RoPE and SwiGLU FFN , improving positional information handling and reconstruction quality. To enable direct flow matching, we align the number of image latent tokens KK in TA-TiTok to match the text encoder’s output sequence length (K=77K=77 for CLIP text encoder).

By integrating these simple yet effective designs across text and image modalities, FlowTok represents both in the same 1D low-dimensional space with shape 77×1677\times 16 (77 tokens, each with 16 dimensions). This compact representation is 3.3×\times smaller than typical 2D flow matching shapes of 32×32×432\times 32\times 4 for image resolutions of 256. This alignment enables fast, direct flow matching and seamless evolution between the two modalities.

Unlike standard flow matching models , FlowTok eliminates the need for intricate conditioning mechanisms, offering a fully self-attention-based generative model. This allows for direct flow across modalities without additional complexity. Unlike CrossFlow , which converts text into 2D embeddings, FlowTok retains the 1D structure of text embeddings, avoiding the need for flattening and transformation into 2D. This simplifies the framework while eliminating reliance on heavy parametric contrastive losses for semantic preservation.

As a result, FlowTok offers a streamlined and resource-efficient training process. Its largest variant, FlowTok-H (1.1B), supports a batch size of 8K on 8 A100 GPUs without requiring gradient checkpointing or gradient accumulation. In contrast, recent text-to-image models of similar scale typically require 32 to 64 A100 GPUs to train with a batch size of only 2K . Moreover, FlowTok converges significantly faster, as shown in Fig. 3(a), with FlowTok-H completing training in just 26.1 8-A100 days—far less than SD 2.1 , which requires 1041.6 8-A100 GPU days.

Inference is also highly efficient, with FlowTok achieving over 10×\times the throughput of Show-o and CrossFlow , as shown in Fig. 3(b). This dramatically reduces computational costs, making text-to-image research far more accessible. To ensure full reproducibility, we train FlowTok exclusively on publicly available datasets, avoiding reliance on high-quality proprietary data. Remarkably, despite its minimalist design and reduced data requirements, FlowTok achieves state-of-the-art text-to-image performance.

Beyond text-to-image generation, FlowTok seamlessly extends to image-to-text generation, maintaining strong performance under the same minimalist framework. We believe our work establishes a strong foundation for future research in generalized cross-modality generation. To support further advancements, all code will be released.

Related Work

Flow Matching. Flow matching models generative processes by constructing a transport map between two distributions via an ordinary differential equation (ODE). It has recently gained traction as the foundation for state-of-the-art text-to-image and text-to-video synthesis models , offering faster training and sampling compared to conventional diffusion methods . Several works further optimize flow trajectories by minimizing curvature . Despite its theoretical flexibility in handling arbitrary distributions, recent approaches primarily evolve noise into target distributions, often relying on complex control signal conditioning, which complicates the pipeline and overlooks the potential of directly transforming control signals into target distributions. In contrast, only a few works explore direct transport within the same modality (e.g., image-to-image ), leaving cross-modal transport (e.g., text-to-image) underexplored. In this work, we introduce FlowTok, a minimal yet effective framework that enables seamless flow across text and image modalities using 1D tokens. Unlike CrossFlow , which follows a similar paradigm but relies on 2D latent representations and incurs additional computational costs due to the text variational encoder, FlowTok operates within a unified, compact 1D token space. This design achieves a 3.3×\times compression rate in latent size, significantly reducing training costs and accelerating the sampling process, all while maintaining state-of-the-art performance.

Text-to-Image Generation. Text-to-image generation has advanced rapidly in recent years, driven by various generative paradigms, including diffusion models , flow matching models , sequence models , and masked generative models . While early works in each category establish the foundation for their respective approaches, subsequent advancements across different model types have primarily emerged from three key areas: careful data collection and advanced image recaptioning for improved data quality , architectural and conditioning improvements for faster convergence and better text-image alignment , and micro-conditioning for finer control over generated samples . By contrast, this work introduces a minimalist framework FlowTok that directly maps text tokens to image tokens, eliminating the need for noise scheduling and complex conditioning mechanisms. This streamlined design enhances both efficiency and simplicity while maintaining competitive performance.

Preliminary

Flow matching is a framework that learns a continuous transformation between a source distribution and a target distribution. The source distribution is not necessarily required to be Gaussian noise, though we use Gaussian noise as a concrete example below.

During training, given a sample XX from the target distribution, a sampled time step t∈t\in, and a noise sample N∼N(0,I)N\sim\mathcal{N}(0,I) from the source distribution. An intermediate representation XtX_{t} is obtained by:

The flow matching model is trained to estimate the velocity field VtV_{t}, which describes the direction from the source to the target distribution. Taking the derivative of XtX_{t} with respect to tt, we have:

where VtV_{t} indicates the direction from the source to the target distribution such that the induced flow accurately transports the source distribution to the target distribution.

Notably, while the source distribution is typically modeled as Gaussian noise in generative frameworks , the flow matching formulation generalizes to arbitrary source distributions, provided that the source and target distributions share the same shape. In FlowTok, we directly define a unified latent space for image and text modalities, treating them as both source and target distributions. This design enables seamless generation across different modalities.

Method

In this section, we focus on text-to-image generation as the primary task to illustrate FlowTok. We first detail how images and text are projected into a unified, compact latent space as 1D tokens while preserving semantic information (Sec. 4.1). Next, we introduce FlowTok as a general framework for seamless flow between text and image tokens and discuss its extension to image-to-text generation under the same formulation (Sec. 4.2).

The structural discrepancy between text and images presents a significant challenge in unifying them within the same latent space for flow matching. Text is inherently semantic, encoded as a 1D latent sequence with high-dimensional channels to preserve meaning, whereas images contain spatially redundant information and are typically represented as 2D feature maps with lower channel dimensions to retain spatial priors. To bridge this gap, we propose encoding images into compact 1D tokens by leveraging recent advancements in image tokenization. This formulation helps preserve the 1D structure of text embeddings, requiring only their projection into a more compressed set of tokens while ensuring that semantic information is retained. Below, we detail how both images and text are encoded.

where T denotes the transpose operation, CE represents the cross-entropy loss, and labels are assigned based on their batch indices, ensuring that each text token is explicitly trained to align with its corresponding CLIP text embedding within the same batch. We also explore alternative approaches to preserving semantic information, such as aligning with the average-pooled text embedding or using a cosine similarity loss with a margin. However, we find that the CLIP-style loss achieves the best performance. More details are provided in Sec. 5.3.

Through the aforementioned designs, FlowTok efficiently tokenizes text into the same low-dimensional latent space while preserving semantic information. This alignment with the tokenized image latent space establishes a foundation for direct flow between compressed text tokens ZT\mathbf{Z}_{\text{T}} and image tokens ZI\mathbf{Z}_{\text{I}}. Notably, when using CLIP as the text encoder, FlowTok effectively reduces the latent size compared to traditional 2D flow matching methods. At an image resolution of 256, the latent size is reduced from 32×32×432\times 32\times 4 to 77×1677\times 16, achieving a 3.3×3.3\times compression. This reduction significantly lowers memory requirements and accelerates training, enhancing the framework’s efficiency and scalability.

2 FlowTok: A General Framework for Seamless Flow Across Text and Image Tokens

Text-to-Image Generation. As shown in Fig. 4 (top), with both image and text mapped into the same latent space, FlowTok leverages vanilla flow matching by stacking DiT blocks . Notably, source modality (i.e., text) is directly treated as the source distribution for flow matching, removing the need for concatenation or cross-attention within the DiT blocks. This design choice further simplifies the overall framework and streamlines text-to-image generation. Combined with the compact 1D tokens introduced in Sec. 4.1, FlowTok achieves high memory efficiency, supporting a batch size of 8K on 8 A100 GPUs. Additionally, it enables fast sampling, running over 10×10\times faster than modern text-to-image diffusion models , significantly lowering the computational barrier for training large-scale text-to-image generative models.

Image-to-Text Generation. FlowTok can also be seamlessly extended to image-to-text generation using compact 1D image and text tokens under the same formulation, as shown in Fig. 4 (bottom). Specifically, the image tokens ZI\mathbf{Z}_{\text{I}} flow to text tokens ZT\mathbf{Z}_{\text{T}}, where a trained text decoder takes ZT\mathbf{Z}_{\text{T}} as input and outputs tokenizer indices, which can then be decoded back into the corresponding caption.

Experimental Results

In this section, we first provide the implementation details of FlowTok (Sec. 5.1), followed by the main results on text-to-image and image-to-text generation (Sec. 5.2). Finally, we present the ablation studies to better understand the design choices of FlowTok in text-to-image generation (Sec. 5.3).

Image Tokenizer. We build our image tokenizer upon the official TA-TiTok codebase with minimal modifications. The encoder uses ViT-B , while the decoder uses ViT-L, both operating with a patch size of f=16f=16. To align with the output sequence length of CLIP’s text encoder, we set the number of 1D latent tokens KK to 77 and the token dimension to 16. Additionally, we enhance the tokenizer with RoPE and SwiGLU FFN . Notably, our enhanced tokenizer achieves a FID of 1.02 in a zero-shot evaluation on the ImageNet validation set, matching the performance of the original TA-TiTok with 128 tokens.

Text Projector. We train a text projector for text-to-image generation, which transforms CLIP text embeddings into a latent representation ZTZ_{\text{T}} of shape 77×1677\times 16, aligning with the image latent space ZIZ_{\text{I}} encoded by our image tokenizer. The text projector consists of six Transformer blocks, each comprising a multi-head self-attention mechanism and a multi-layer perceptron (MLP), both enhanced with skip connections to ensure stable training.

Text Decoder. We train a text decoder for image-to-text generation, composed of six Transformer blocks, similar to the text projector. The decoder takes the text latent representation ZTZ_{\text{T}} as input and outputs the corresponding CLIP text tokenizer indices, which can be further converted into text using the CLIP text tokenizer.

FlowTok. We adopt DiT blocks as the fundamental building units of our FlowTok to model token interactions. Specifically, we follow the DiT architecture to implement FlowTok-B for efficient ablation studies and FlowTok-XL for enhanced performance. To further push the performance, we scale up the depth, width, and number of attention heads, constructing FlowTok-H with 1.1B parameters. The detailed model configurations are provided in Tab. 1.

Dataset. We employ open-source datasets to facilitate the reproducibility of FlowTok’s simple framework. Specifically, our image tokenizer is trained on DataComp-1B , and our text tokenizer is trained on COCO . For text-to-image generation, inspired by recent works , we adopt a two-stage training strategy: pre-training and fine-tuning. The pre-training stage leverages a combination of DataComp-1B , CC12M , and LAION-aesthetic , while the fine-tuning stage incorporates additional high-quality datasets, including LAION-art , LAION-pop , JourneyDB , and DALLE3-1M . For image-to-text generation, we follow the Karpathy split of COCO to divide the training and validation sets. Detailed dataset information is provided in the Appendix.

Training. The training objectives of FlowTok primarily focus on predicting velocity in flow matching, denoted as Lfm\mathcal{L}_{\text{fm}}. For text-to-image generation, we introduce two additional losses: KL-divergence loss (Lkld\mathcal{L}_{\text{kld}}) to enforce a Gaussian distribution on text tokens and text alignment loss (Lalign\mathcal{L}_{\text{align}}) to preserve semantic information as discussed in Sec. 4.1. Formally, the overall training objective is:

where γ1\gamma_{1} and γ2\gamma_{2} control the weighting of losses. By default, we set γ1\gamma_{1} to 1×10−41\times 10^{-4} and γ2\gamma_{2} to 11 for text-to-image generation, while both are set to 0 for image-to-text generation.

Evaluation. We follow standard evaluation practices to report relevant metrics for both text-to-image and image-to-text generation. Specifically, for text-to-image generation, we report FID-30K on the COCO , and FID on MJHQ-30K . For image-to-text generation, we report BLEU-4 , METEOR , ROUGE , CIDEr , and SPICE on the COCO Karpathy Split . To incorporate classifier-free guidance (CFG) within FlowTok, we follow CrossFlow and utilize a CFG indicator. Unless otherwise stated, we find that using only 20 steps for sampling is sufficient due to the small 1D latent shape of FlowTok. This significantly speeds up the inference process, enabling faster generation without compromising performance.

2 Main Results

Text-to-Image Generation. We report zero-shot text-to-image generation results on COCO and MJHQ-30K in Tab. 2. The compared methods are categorized into two groups: text as conditions, where text serves as a guiding signal for image generation, and text as source distributions, where text is directly modeled as a distribution in the generative process. As observed, FlowTok achieves comparable performance to prior methods in both categories on COCO FID-30K. Specifically, compared to CrossFlow , which also uses text as the source distribution, FlowTok-H attains a FID-30K of 9.67—roughly on par with CrossFlow. When further evaluating FlowTok on MJHQ-30K to assess the aesthetic quality of generated images, we find that, despite being trained solely on publicly available datasets without access to high-quality proprietary data, FlowTok-XL already surpasses other state-of-the-art models, demonstrating its ability to generate diverse, high-quality images. Furthermore, FlowTok-H further improves the FID score to 7.15, underscoring its superior image generation capabilities.

Beyond performance, FlowTok requires significantly fewer training resources compared to existing state-of-the-art models. Specifically, FlowTok-XL completes training in just 20.4 8-A100 days, while FlowTok-H increases the budget slightly to 26.1 8-A100 days. In contrast, the most efficient text-as-condition model, PixArt-α\alpha , still demands 94.1 8-A100 days. Compared to CrossFlow , which also treats text as source distributions and requires 78.8 8-A100 days, FlowTok is much more efficient.

Additionally, FlowTok demonstrates significantly faster inference speeds. At 256px resolution, FlowTok-XL generates 22.7 images per second, while FlowTok-H achieves 18.2 images per second. In contrast, PixArt-α\alpha runs at 7.9 images per second, and Show-o at just 1.0 images per second. More notably, within the text as source distributions category, FlowTok achieves a 20×\times speedup in sampling time compared to CrossFlow, which runs at only 1.1 images per second. This efficiency stems from FlowTok ’s streamlined framework and its effective use of 1D tokens, significantly reducing computational overhead.

Image-to-Text Generation. We evaluate image-to-text generation on COCO using the Karpathy split , with results summarized in Tab. 3. To ensure a fair comparison, we categorize methods into two groups: direct flow from image to text distributions, which represents a new paradigm leveraging flow matching for direct image-to-text transformation, and other methods, considering only those trained without CIDEr optimization. Within the direct flow category, FlowTok-XL consistently outperforms its counterpart, CrossFlow , across most metrics. Specifically, FlowTok-XL achieves a BLEU-4 (B@4) score of 37.1, surpassing CrossFlow by 0.7, and a CIDEr score of 117.0, exceeding CrossFlow by 0.8. Moreover, compared to state-of-the-art methods from other paradigms, FlowTok-XL demonstrates competitive performance, highlighting direct flow matching as a promising approach for image-to-text generation. Notably, FlowTok performs image-to-text generation under the same formulation using compact 1D tokens, theoretically requiring fewer training resources and enabling faster sampling compared to paradigms that operate on 2D latents, as adopted by CrossFlow. However, a direct quantitative comparison is not possible, as CrossFlow has not released the corresponding checkpoint for evaluation.

3 Ablation Studies

We conduct ablation studies on text-to-image generation using FlowTok-B and evaluate them on COCO for efficiency. Our ablations focus on the design of the text alignment loss, as it plays a critical role in preserving semantic information. Specifically, we investigate three key aspects: the text alignment target (Tab. LABEL:tab:ablation_loss_target), the choice of loss function (Tab. LABEL:tab:ablation_loss_function), and the loss weight (Tab. LABEL:tab:ablation_loss_weight). Details are provided below.

Text Alignment Target. We first investigate the choice of alignment target for the projected text tokens ZT\mathbf{Z}_{\text{T}} in Tab. LABEL:tab:ablation_loss_target. A straightforward baseline is to directly apply average pooling (row 1) to the original CLIP text embedding Tinit\mathbf{T}_{\text{init}} along the channel dimension, reducing the dimensionality from 768 to 16 to match ZT\mathbf{Z}_{\text{T}}. However, this approach performs significantly worse compared to using a simple MLP to learn the alignment target (row 2), as inspired by prior works . We attribute this performance gap to the fact that adjacent channels in the CLIP text embedding are not necessarily correlated, and simple average pooling discards too much semantic information. In contrast, a learnable MLP mitigates this information loss, making it a more effective choice for defining the text alignment target.

Text Alignment Loss Function. Next, we examine the text alignment loss function in Tab. LABEL:tab:ablation_loss_function. Besides the contrastive loss adopted from , we explore using a cosine similarity loss, similar to . Specifically, we compute the cosine similarity between the text tokens and the alignment target, applying a penalty to pairs with similarity below a threshold. Our experiments show that while both loss functions are effective, the contrastive loss achieves better performance.

Text Alignment Loss Weight. Finally, we investigate the impact of the text alignment loss weight, γ2\gamma_{2} in Tab. LABEL:tab:ablation_loss_weight. Our results indicate that setting γ2\gamma_{2} to 1.0, equal to the weight of the flow matching loss, is sufficient to preserve semantic information while maintaining high-quality image generation. Increasing γ2\gamma_{2} further can cause the text alignment loss to dominate the overall objective during early training stages, potentially hindering final performance.

Conclusion

In this paper, we introduce FlowTok, a minimal yet powerful framework that enables seamless direct flow between 1D text and image tokens. Through carefully designed key modules and loss functions, FlowTok projects both modalities into a unified 1D latent space while preserving semantic information, enabling both text-to-image and image-to-text generation under the same formulation. This design makes FlowTok highly memory-efficient, supporting an 8K batch size on just 8 A100 GPUs during training. Additionally, its simplicity accelerates convergence—within approximately 20 days on 8 A100 GPUs, FlowTok achieves performance comparable to state-of-the-art models that require significantly longer training times. The streamlined design also enables over 10×\times faster sampling than modern text-to-image generative models. By releasing our code, we aim to further advance research in text-image cross-modal generation.

Appendix

The supplementary material provides additional information:

Sec. A: More implementation details, including dataset filtering, and FlowTok training hyperparameters.

Sec. B: Additional qualitative text-to-image and image-to-text generation samples produced by FlowTok.

Sec. C: Discussions on limitations and future work of FlowTok.

A More Implementation Details

Dataset Filtering. In line with previous works , we apply three filtering criteria to curate open data for training the image tokenizer and FlowTok: resolution, aesthetic quality, and watermark filtering. The COCO dataset is used directly to train the text decoder without any filtering. Details of the applied filtering criteria are shown in Tab. 5.

Specifically, resolution filtering is applied during the training of the image tokenizer and for text-to-image generation. This ensures that the longer side of each image is at least 256 pixels and the aspect ratio is below 2. For text-to-image training, we further apply aesthetic filtering using the LAION-aesthetic-v2 predictor to retain only high-quality images. Images with aesthetic scores above 5.0 are retained during the pre-training stage, while a stricter threshold of 6.0 is used during fine-tuning to ensure even higher image quality.

Additionally, watermark filtering is implemented for FlowTok’s text-to-image generation by using the LAION-WatermarkDetector , removing images with watermark probabilities exceeding 0.5. Synthetic datasets such as JourneyDB and DALLE3-1M are exempt from these filtering steps, as they inherently meet our high resolution and quality standards.

Training Hyper-parameters. Tab. 6 provides the complete list of hyper-parameters used for training FlowTok.

B Qualitative Examples of FlowTok

Additional Generation Results. Fig. 5, Fig. 6, and Fig. 7 present additional text-to-image generation results produced by FlowTok, demonstrating its ability to generate diverse, high-fidelity images. Meanwhile, Fig. 8 displays the image-to-text generation results, showcasing FlowTok’s capability to produce accurate and descriptive captions.

C Limitations and Future Work

The primary limitation of FlowTok arises during text-to-image generation. To match the compact dimensionality of image latents (e.g., 16), FlowTok projects CLIP text embeddings into the same low-dimensional latent space. While the text alignment loss helps preserve semantic information, some degree of information loss is inevitable during this projection. Consequently, the alignment between text and generated images may be weaker compared to state-of-the-art models employing cross-attention mechanisms. To address this, one potential solution is to introduce a stronger alignment loss that better retains textual semantics. A more fundamental approach, however, involves increasing the channel dimensionality of image latents by aligning them with vision foundation models during image tokenizer training, as suggested by VA-VAE . This strategy aims to identify an optimal channel dimension for both image and text tokens, achieving a balance between preserving semantic information and maintaining efficiency in training and inference.

Additionally, FlowTok currently utilizes only the vanilla flow matching technique to validate the framework’s effectiveness. However, many recent advancements in flow matching, such as logit-normal sampling , have not yet been explored in our model. Incorporating these techniques could accelerate convergence and enhance performance.

Finally, FlowTok serves as a starting point for exploring efficient direct evolution between text and image modalities. In the future, we aim to extend to a more general framework that can accommodate a broader range of modalities, supporting additional tasks under the same unified formulation.

References