ReasonGen-R1: CoT for Autoregressive Image generation models through SFT and RL

Yu Zhang, Yunqi Li, Yifan Yang, Rui Wang, Yuqing Yang, Dai Qi, Jianmin Bao, Dongdong Chen, Chong Luo, Lili Qiu

Introduction

Recent breakthroughs in models such as OpenAI’s o1 and Deepseek’s R1 have demonstrated the significant advantages of reinforcement learning (RL) methods for enhancing the thinking and reasoning capabilities of large language models (LLMs). These advancements confirm that step-by-step reasoning substantially improves answer accuracy and robustness. Naturally, transferring the powerful reasoning capabilities of LLMs from text-based tasks to image generation tasks becomes critically important. Models such as ChatGPT , Gemini , and Janus-Pro have introduced unified image-generation paradigms, highlighting that multimodal LLM-based autoregressive content generation exhibits superior instruction-following abilities and image quality. Consequently, a crucial next step is exploring how to effectively incorporate thinking and reasoning via RL into autoregressive generation models.

Motivated by human creative processes, where artists typically contemplate structural and sequential considerations before image creation, we aim to enable autoregressive generation models to autonomously produce textual reasoning sequences based on user prompts. This approach harnesses the inherent instruction-following strengths of autoregressive image generation models, encouraging them to make creative associations and structural decisions akin to human thought processes.

However, achieving this goal introduces several challenges. First, current autoregressive image-generation models, such as Janus-Pro, typically generate images directly from text prompts without the ability to concurrently produce textual reasoning. Consequently, straightforward RL supervision might be ineffective. Second, designing an efficient RL-based post-training pipeline that facilitates "thinking-based" generation within autoregressive image-generation frameworks remains unexplored and unvalidated.

To address these challenges, we introduce ReasonGen-R1, a novel two-stage training paradigm combining supervised fine-tuning (SFT) with chain-of-thought (CoT) and RL using Group Relative Policy Optimization (GRPO) , tailored explicitly for pretrained autoregressive image-generation models. To overcome the first challenge, we employ an "instruction → CoT → image" pipeline to jointly supervise textual reasoning sequences and image outputs during SFT training. We constructed a comprehensive dataset comprising 200k samples from the LAION aesthetics subset , meticulously annotated using GPT-4.1-mini to include rich CoT reasoning trajectories for each triplet, covering diverse reasoning scenarios. This dataset effectively balances the dual objectives of CoT text generation and target image generation without compromising image quality.

For the second challenge, ReasonGen-R1 employs an efficient GRPO framework utilizing the powerful image-understanding model Qwen-2.5-VL as the reward model. For each training rollout, we assess prompt–image alignment by querying Qwen-2.5-VL for a binary consistency score. Additionally, we found Reinforcement Learning with interleaved modality output extremely sensitive to entropy explosion or entropy vanishing. To tackle this, we introduce a unique adaptive entropy loss design to ensure stable and effective training.

Empirical results in Table 2 demonstrate that ReasonGen-R1 achieves superior performance on benchmark datasets such as GenEval (+6%), DPG-Bench (+1.69%) and T2I-Benchmark (+13.38%), significantly enhancing reasoning-based generation capabilities.

We are the first to integrate the reasoning process into autoregressive image generation via a two-stage SFT + GRPO training framework, establishing a strong baseline for future “think-and-generate” content creation.

Technically, we build a large-scale CoT image-generation dataset to spark the model’s exploration of reasoning, then leverage multimodal LLMs as reward models within GRPO to guide the model toward more suitable reasoning for I2T tasks; an adaptive entropy loss further mitigates entropy explosion or collapse.

Extensive experiments confirm the efficiency and effectiveness of the proposed ReasonGen-R1 framework across multiple benchmarks.

Related Work

Recent work on unified auto-regressive generation models has demonstrated that a single model can both generate language and images. By mapping images into the same embedding space as text and feeding both into an LLM backbone, these systems attain a richer understanding of prompts and can produce multimodal outputs. However, most existing designs—such as Janus-Pro —still generate text and images in distinct phases by default. This modality split prevents true interleaving of words and pixels in one continuous sequence, limiting their effectiveness on tasks that demand integrated, cross-modal reasoning.

2 Chain-of-Thought in LLM

Chain-of-thought (CoT) reasoning has emerged as an effective strategy in large language models (LLMs), allowing models to decompose complex tasks into intermediate logical steps. This technique has led to state-of-the-art results in mathematical problem-solving, commonsense reasoning, and compositional tasks in models such as PaLM , GPT-4 , and LLaMA 2 . Wei et al. and subsequent works have shown that reasoning traces not only improve model performance but also enable interpretability and controllability.

Our approach builds upon this line of work by tightly integrating CoT reasoning into the autoregressive image generation process. We train the model to generate reasoning rationales and visual tokens within a single coherent sequence, enabling end-to-end optimization for both interpretability and performance.

3 Reinforcement Learning in LLM and LVLM reasoning

More recently, reinforcement learning (RL) has been employed to improve the reasoning capabilities of LLMs and vision-language models (VLMs). For instance, leading by DeepSeek-R1 , a number of works utilizes Group Relative Policy Optimization (GRPO) and its variants, to enhance its reasoning ability without the need for a separate critic model. This approach normalizes rewards within a group of generated outputs, reducing computational cost and improving performance.

Despite reinforcement learning (RL) has been increasingly applied to refine image generation models, particularly in text-to-image synthesis , the integration of RL specifically for reasoning within image generation remains largely unexplored. In this work, we explore the possibility of the magical reasoning ability for enhanced image generation.

Method

ReasonGen-R1 consists of two main parts: (1) supervised finetuning on a base autoregressive image generation model, equipping the model with textual reasoning ability. (2) reinforcement learning on the finetuned model, further enhancing the model’s capability to analyze the prompt and output the final image.

To equip the base model with the ability to generate interleaved text-image outputs, we begin by constructing a diverse dataset consisting of short prompts, dense prompts, and corresponding images. Specifically, we first select 200,000 images from the LAION-Aesthetic V1 dataset, which is a subset of the LAION-5B collection . Since the base model, Janus-Pro, is restricted to generating square images, we crop the long side of each image to match the shorter side, resulting in a square output.

Next, we query GPT-4.1 mini to generate a concise caption for each image, focusing on key details such as object color, counts, spatial relationships, and other contextual elements. Then, we use GPT-4.1 nano to augment the concise caption, generating additional prompts to increase diversity. These augmented prompts include a set of image tags, object-centric phrases, three paraphrased versions of the concise caption, and one varied caption written in a different style. At the same time, we also generate a detailed caption, providing a longer, more comprehensive description of the image. Notably, the concise caption is generated directly from the image input, while the augmented prompts and detailed caption are generated solely from the concise caption. This approach ensures that GPT-4.1 does not introduce additional information from the image, preventing potential information discrepancies during the SFT training as the generation model being trained only has access to the concise caption.

The result is a high-quality, diverse dataset consisting of concise captions paired with augmented concise captions and detailed captions.

1.2 SFT Training

Enabling the model to interleave textual reasoning with image generation presents a core challenge: most auto-regressive generators currently produce text and images in separate sequences, requiring explicit user instructions to switch modalities.

In our supervised fine-tuning (SFT) stage, we address this limitation by training the model to first generate a coherent reasoning rationale and then seamlessly transition to image synthesis within a single sequence. We build on Janus-Pro-7B as the base model, which uses a special image-start token to trigger visual output. As illustrated in Figure 1, each training sequence begins with a concise image prompt followed by the detailed reasoning caption. We then insert the image-start token and the corresponding image tokens. Through this formatting, the model learns to produce a detailed rationale before autonomously emitting the image-start token and generating the final image.

2 Reinforcement Learning

Group Relative Policy Optimization (GRPO) has shown a strong capability to explore the reasoning potential of LLM models. To further align generated images with the text-based rationale and input prompt, we adapt Group Relative Proximal Optimization (GRPO) for image generation.

The GRPO policy objective then mirrors PPO’s clipped surrogate, plus a KL penalty:

Finally, each per-token importance weight is

Beyond GRPO, our RL algorithm adds an adaptive entropy loss to further enhance the stability of the training.

Adaptive entropy loss is inspired by SAC , a RL algorithm that is commonly used in Robotics RL. More specifically, adaptive entropy loss sets a target entropy. And during training, the entropy loss coefficient would automatically update through gradient descent. The adaptive entropy loss function in is:

where the α\alpha is the learnable entropy loss term, Htarget\mathcal{H}_{\text{target}} stands for the target entropy and log⁡π(at∣st)\log\pi(a_{t}|s_{t}) stands for the entropy of the current action output conditioned on the current state in robot learning. The learnable entropy loss term α\alpha, often parameterized as log⁡(ϕ)\log(\phi), is always positive, encouraging the model to explore more.

In our setting, we found that the model is very prone to both entropy vanishing and entropy explosion, often leading to mode collapse in image generation. To tackle this, we modify the parameterization method of α\alpha to arcsin⁡(ϕ)\arcsin(\phi), so that we can both learn positive and negative α\alpha. The complete loss term for updating α\alpha is:

In addition, our RL algorithm uses batch sub-sampling to remove groups that score all 0 or all 1 and remove the KL loss, followed by DAPO .

Our final objective function is therefore written as:

2.2 Reward Design

As opposed to rule-based reward in original GRPO in LLM training, it’s difficult to have a golden pre-defined rules to assess the consistency between prompt and generated image. To tackle this, we use a strong VLM, Qwen-2.5-VL 7B , as our reward model. In each rollout, the generation model produces a single sequence: prompt → reasoning → image. We compute rewards solely on the generated image by querying a pretrained vision–language model (VLM) to evaluate consistency between the input text and output image. To credit the preceding reasoning steps, we propagate the image-level reward back through the entire sequence, reinforcing textual rationales that yield higher-quality visual outputs.

Experiment

In our experiment section, we aim to answer the following questions:(1) To what extent does incorporating textual reasoning improve instruction adherence in image generation? (2) How much does the RL benefit from the SFT training warmup? (2) How much does the RL training benefit from the size of reward model? (4) How does the adaptive entropy loss benefit the training?

T2I Benchmark : 6,000 compositional prompts spanning attribute binding, object relations, and complex scene layouts (color, shape, texture bindings; spatial and non-spatial relations).

GenEval : Object-focused alignment tasks designed to assess fine-grained consistency between text and image outputs.

DPG-Bench : Dense-prompt generation emphasizing detailed, descriptive instructions.

These benchmarks collectively cover a wide spectrum of compositional and alignment challenges, providing a thorough evaluation of the model’s ability to follow complex textual directives.

1.2 Dataset Settings

Our training dataset comprises prompts drawn from three benchmarks: GenEval, DPG-Bench, and T2I-CompBench++ . For GenEval, we enlarge its object vocabulary from 80 to 308 and extend its original generator with a new variant that specifies two distinct objects along with their respective counts. Using this augmented generator, we synthesized 12,552 unique prompts and filtered out any that overlap with the GenEval test set, resulting in 12367 prompts. For the DPG-Bench, we leveraged GPT-4.1 to produce 5,000 fresh prompts: for each draft, we sampled five existing DPG prompts at random and instructed GPT-4.1 to craft a final prompt matching their length and style. Lastly, we incorporated all prompts from the official T2I-CompBench++ training split without modification, resulting in 11,003 prompts.

1.3 Training Settings

In all supervised fine-tuning (SFT) experiments, we trained the model for 1 epoch and selected the final checkpoint. For reinforcement learning (RL) experiments, we trained the model for 300 steps and chose the checkpoint with the highest validation reward.

2 Main Result

To answer the first question: To what extent does incorporating textual reasoning improve instruction adherence in image generation? We compare ReasonGen-R1 against the Janus-Pro-7B baseline and leading diffusion and auto-regressive text-to-image systems across three diverse benchmark suites. As shown in Table 1, 2, 3 ReasonGen-R1 outperforms the base model in all three benchmarks. ReasonGen-R1 also surpasses many leading image-generation models. These results indicate that the textual-reasoning model framework and SFT-RL pipeline greatly boost the performance of the auto-regressive image generation model.

3 Ablation Study

To answer the second research question: How much does the RL benefit from the SFT training warmup? We compared ReasonGen-R1 with pure RL. Since the original model doesn’t have the capability to control its output modality during RL rollout, we first instruct the model to output reasoning text. Once it finishes text generation, we swap its last end of sentence token with an image start token to start image generation.

As shown in Table 4, ReasonGen-R1 significantly outperforms the w/o SFT baseline by 18%, indicating that SFT primes the base model to carry out proper interleaved reasoning and generation. The w/o RL variant, however, reveals that SFT alone is not sufficient, because the GPT-annotated CoT traces do not always represent the reasoning trajectories most conducive to high-quality image synthesis. Nevertheless, the gap between w/o SFT and w/o RL shows that SFT equips the model to explore diverse thinking paths; the subsequent RL stage is therefore essential to unlock this potential, fully realize the “think-and-generate” motivation, and achieve the final performance gains.

3.2 Reward Model Size Matters

For our third research question, we use Qwen-2.5-VL-3B as the rewarder VLM for comparison. As shown in Table 4, a smaller VLM failed to provide good reward signals, leading to poor performance after RL. This indicates that using a large and accurate rewarder model is crucial to our RL algorithm.

3.3 Stable Training with Adaptive Entropy Loss

To address our final research question, we perform a comparative analysis of our model under two configurations. The first configuration involves RL training the model without any entropy loss, allowing the model to learn without any explicit regularization on entropy. The second configuration incorporates a fixed entropy loss term, where the entropy of the model’s predictions is penalized by a constant value during training. This setup allows us to evaluate the impact of adaptive entropy regularization on the model’s performance and the stability of RL training.

As illustrated in Figure 3, RL without entropy loss experiences entropy explosion after 100 training steps, resulting in degraded performance. On the other hand, applying a fixed entropy penalty of -0.002 causes a continuous decline in entropy, reaching dangerously low levels at step 80, which leads to mode collapse and a sharp decrease in reward. These observations highlight the challenges of RL training with interleaved text and images, particularly its sensitivity to entropy loss regularization. In contrast, our adaptive entropy loss effectively maintains entropy within an optimal range, ensuring stable training.

Conclusion

In this paper, we introduce ReasonGen-R1, a novel two-stage framework combining Chain-of-Thought (CoT) reasoning with reinforcement learning (RL) for improving autoregressive image generation. Our approach addresses the challenges of interleaving textual reasoning with image generation by employing supervised fine-tuning (SFT) followed by reinforcement learning using Group Relative Policy Optimization (GRPO). We show that incorporating CoT reasoning into the image generation process results in better adherence to complex instructions, leading to significant improvements in image quality and consistency with textual prompts.

Through extensive experimentation on benchmark datasets such as GenEval, DPG-Bench, and T2I-Benchmark, we demonstrate that ReasonGen-R1 outperforms previous models, including Janus-Pro and other state-of-the-art systems, across multiple evaluation metrics. Additionally, our results emphasize the importance of a well-calibrated reward model and adaptive entropy loss in ensuring stable and effective RL training.

The proposed framework offers a promising direction for advancing multimodal generative models by integrating textual reasoning capabilities into autoregressive image generation tasks. Our work lays the foundation for future research exploring further improvements in reasoning-based generation, multi-step planning, and fine-grained control over image content generation. We plan to release our dataset and training code to facilitate the continued development of this area of research.

Limitation

Despite the promising results observed in this study, there are several limitations that need to be addressed in future work. First, our approach primarily focuses on a specific set of benchmarks, which may not fully represent the diversity of real-world tasks. While the results on these benchmarks indicate great improvements, the generalization of the model to more complex or domain-specific tasks remains to be further investigated.

Second, the reliance on large-scale pretrained models, such as GPT-4.1, introduces potential biases stemming from the data used in SFT stage. These biases may affect the robustness and fairness of the generated outputs, particularly in sensitive or underrepresented contexts.

Finally, while we have implemented adaptive entropy loss to mitigate mode collapse, the sensitivity of this parameter to the specific task needs to be better understood.

References

Appendix A Method Details

Our dataset construction involves three distinct calls to the OpenAI API, each serving a different role:

We use GPT-4.1 Small to generate a short, accurate, and informative caption that highlights object counts, colors, positions, and other details. During this stage, we feed the GPT with image.

You are a data annotation expert. Generate a concise image caption for the image, focusing specifically on the color, number, position, and other details of objects, background, and humans present. Analyze carefully and ensure accuracy in your description. The caption should be a single short sentence that faithfully includes most of the important information in the original caption.

We use GPT-4.1 Nano to augment the concise caption obtained from the previous API call. We augment the caption into several categories. The purpose for the augmentation is to prevent the model from overfitting to one specific prompt pattern during the SFT training. During this stage, we don’t give the image to the GPT, concise caption from the previous call would be the only input.

You are an image-annotation augmentation expert. Given the original detailed caption below, analyze it and produce a single JSON object (and only that JSON) with the following top-level keys—no nested structures: • “concise_caption”: A one-sentence compressed caption capturing the main objects, their colors and positions. • “paraphrases”: An array of 3 alternative one-sentence phrasings that preserve every key details but vary word order and synonyms. • “tags”: An array of 5–8 keywords describing objects, colors, positions, and scene. • “varied_captions”: An array of 3 one-sentence captions, each in a different style of your choice (the model should decide the styles randomly). • “object_prompts”: An array of several very short prompts in the form “a/an/number optional adj. n.”, listing only the main object noun (no more than 3) in descending order of their significance.(e.g. “a clock”, “two wooden chairs”). **Input (detailed caption):** “{CONCISE_CAPTION}” **Expected Output (json):** { "concise_caption": "...", "paraphrases": ["...", "...", "..."], "tags": ["...", "...", "...", "...", "..."], "varied_captions": ["...", "...", "..."], "object_prompts": ["...", "..."] }

We use GPT-4.1 Nano to generate a detailed caption from each concise caption. This detailed caption serves as the ground truth chain-of-thought (CoT) supervision during the supervised fine-tuning (SFT) stage, guiding the model to learn reasoning based solely on the concise prompt. Importantly, GPT is provided only with the concise caption and not the corresponding image when generating the detailed caption. This design choice ensures that no additional visual information leaks into the supervision, preventing an information gap during SFT—where the model only has access to the concise caption. Without this precaution, the model may learn to generate overly imaginative or irrelevant reasoning that is not grounded in the available input.

You are an image-annotation augmentation expert. Given the following inputs: Input “concise_caption”: A concise description of the image (e.g. “A red clock on a wooden table”) • “expanded_prompt”: A richly detailed prompt that (1) restates the concise_caption with full color, count, position, background, and mood. **Inputs**: concise_caption = "{concise_caption}" **Output** (Please directly output the expanded_prompt):

A.2 SFT Training Details

Prompt Augmentation To prevent overfitting to a single prompt format, we apply prompt augmentation strategies described in Appendix 2. Concise Image Caption Augmentation to enhance prompt diversity. Specifically, during training, we uniformly sample one augmentation type (treating the original concise caption as one of the types) to replace the original prompt. If the selected type is tags or object_prompts, we concatenate all items using commas. If the type is paraphrases or varied_captions, we randomly select one candidate from the list. If concise_caption is selected, we use it directly without modification.

Prompt Formating To reduce the distribution gap between the base model and our target model, we add a bridging prompt between the concise prompt and the ground truth CoT. More specifically, we add the following:

A.3 RL Algorithm Details

Adaptive Entropy Loss To stabilize reinforcement learning with interleaved text and image outputs, we employ an Adaptive Entropy Loss . We have two independent Adaptive Entropy Loss regularizers for each modality. Because image and text tokens have vastly different vocabulary sizes—and consequently different natural entropy scales—we maintain separate entropy targets for each. Specifically, we use a target entropy of 7.0 for image tokens and 2.0 for text tokens. These target entropy values come from the average rollout entropy observed after supervised fine-tuning.

Reward Model We use Qwen-2.5-VL-7B as our reward VLM model for reinforcement learning. During training, we’ll prompt it with the following template for each rollout image:

You are given a text prompt: "{prompt}" Below is one generated image: 1. Describe the image thoroughly (objects, colors, layout, etc.), do not be affected by the prompt. 2. Identify key visual elements and instructions from the prompt. 3. Evaluate how well the image follows the prompt: - Are all required elements present? - Are object counts, colors, and positions accurate? Be extremly strict and precise: Only if the image matches the prompt perfectly, respond with: \boxed{1}. Otherwise, respond with: \boxed{0} Reason before your final boxed answer. Only one number should appear inside the box.

Other Implementation Details Our reinforcement learning framework builds on verl , a flexible, efficient, and production-ready library for training large language models. By leveraging verl, we streamline our RL pipeline and maximize training throughput.

During RL, we generate images using a classifier-free guidance scale of 1.0. We found this is sufficient to generate meaningful images for the reward VLM model to grade and it can greatly accelerate the RL rollout speed as we don’t need to generate the unconditioned images. During inference and evaluation, we use a classifier-free guidance scale of 5.0, the same as the default value of our base model Janus-Pro 7B.

Appendix B Experiment Details

Visual Sensitivity to Chain-of-Thought Substitutions To further examine whether the chain-of-thought (CoT) genuinely guides the image generation process, we conduct a controlled substitution analysis. For each example, we selectively alter or add a specific element within the original CoT—such as object property, lighting condition, or background setting—while keeping the prompt and the rest of the reasoning unchanged. The goal is to isolate the impact of the modified token or phrase and observe how it propagates through the model’s internal planning and ultimately manifests in the generated image.

As illustrated in Table LABEL:fig:cot_case_vis, these targeted CoT substitutions lead to consistent and interpretable changes in the visual output. For instance, replacing “warm sunlight glows softly” with “bright sunlight shines” results in a stark contrast in overall lighting tone and shadow definition.

These results provide strong qualitative evidence that the model’s generation is causally entangled with its reasoning process. The images do not merely correlate with the CoT—they reflect a coherent execution of its planning steps. This controlled substitution strategy thus offers compelling support for the claim that ReasonGen-R1 uses its CoT to explicitly anchor and shape each scene’s content and style.

Chain-of-Thought Word Frequency Analysis Figure 5 exposes a clear pattern in ReasonGen-R1’s chain-of-thought. First, it anchors each scene with high-level framing—“sense,” “scene” and “natural” dominate, appearing in over 140% of CoTs—emphasizing overall context and realistic setting. Then, it refines visual style: terms like “soft,” “highlights,” “mood,” and “sleek” (all above 100 %) specify lighting quality, emotional tone and texture.

Critically, the presence of “highlighting” and “emphasizing” (each in at least 70 % of CoTs) signals an explicit step to draw attention to the main subject. This reveals that ReasonGen-R1 doesn’t merely describe objects; it actively plans compositional focus.

In addition to its core lexicon, ReasonGen-R1 draws on a sprawling array of less frequent modifiers—“background,” to establish environmental context; “features,” to spotlight salient visual elements; “calm,” to evoke a serene atmosphere; “moments,” to impart a sense of temporal capture; “captured,” to underscore photographic realism; and many more—to infuse each reasoning sequence with subtle, context-specific nuance.

Overall, this analysis shows that ReasonGen-R1’s chain-of-thought leverages complementary components—scene framing, style detailing, subject highlighting, and narrative enrichment—in concert to guide image generation.

B.2 Experiment Settings

Table 6 summarizes the key settings used during supervised fine-tuning (SFT) and reinforcement learning (RL). These choices balance model capacity, sequence coverage, and compute efficiency for each stage of our pipeline.

B.3 Case Visualization

In this section, we present visualizations of ReasonGen-R1’s rollouts across four prompt categories: long and detailed prompts (Table LABEL:fig:long_case_vis), counting prompts (Table LABEL:fig:count_case_vis), spatial-relationship prompts (Table LABEL:fig:spatial_case_vis), and complex-attributes prompts (Table LABEL:fig:complex_case_vis). The prompts in the first category are sampled from DPG-bench, while those in the remaining three categories are generated using Geneval.

Across all categories, the chain-of-thought (CoT) produced by ReasonGen-R1 aligns closely with the content of the generated images, demonstrating that its internal “thinking” effectively guides the planning and composition of each scene.

When using chain-of-thought (CoT) generation, ReasonGen-R1 consistently transforms terse prompts into richly detailed descriptions by specifying object attributes, lighting, textures, backgrounds, and overall mood. For instance, given the prompt “A photo of two persons” (Figure LABEL:fig:two_person), ReasonGen-R1 first refines “two persons” into “a young woman and a man,” then adds “natural light” and clothing details, follows with an ambient background description, and finally articulates the desired emotional tone of the scene.