Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization

Weiyun Wang, Zhe Chen, Wenhai Wang, Yue Cao, Yangzhou Liu, Zhangwei Gao, Jinguo Zhu, Xizhou Zhu, Lewei Lu, Yu Qiao, Jifeng Dai

Introduction

With the remarkable success of large language models (LLMs) in the field of natural language processing, the training paradigm comprising pre-training and supervised fine-tuning (SFT) have also swept the multimodal field, becoming the primary choice for the research and development of multimodal large language models (MLLMs). Benefiting from the large-scale pre-training corpora and high-quality SFT data , a series of open-source MLLMs exhibit strong performance across various domain and tasks, some even achieving results comparable to commercial models such as GPT-4o and Gemini .

However, open-source MLLMs still exhibit limited reasoning capabilities. As shown in Figure 1, InternVL2-8B achieves a score of 58.3 on MathVista , a benchmark for multimodal reasoning, when using direct answers but drops to 56.8 with Chain-of-Thought (CoT) reasoning, indicating that CoT reasoning actually reduces its performance. This decline is commonly observed across open-source MLLMs . We attribute this phenomenon primarily to a distribution shift introduced by the SFT loss. Specifically, SFT relies on teacher forcing, where the model is trained to predict the next token based on previous ground-truth tokens. However, during inference, models must predict each token based on their own prior outputs, leading to a distribution shift between training and inference. Since the direct-answer approach requires only brief responses, while CoT reasoning involves generating a long rationale, the distribution shift problem becomes more severe during CoT. This results in models performing worse with CoT reasoning compared to direct-answer responses.

To address the limitations of CoT reasoning in MLLMs, we draw inspiration from recent NLP approaches that use Preference Optimization (PO) techniques to align model outputs with desired reasoning patterns. Specifically, methods like Direct Preference Optimization (DPO) allow models to learn from preference signals to generate responses that better align with user requirements, offering the foundation for Reinforcement Learning from Human Feedback (RLHF). While RLHF has been explored for MLLMs primarily to reduce hallucinations , its application for enhancing multimodal reasoning remains under-explored. Building on these insights, we conduct a systematic study on using PO to strengthen the multimodal reasoning capabilities of MLLMs.

Enhancing the multimodal reasoning abilities of MLLMs through PO presents several challenges: (1) Limited multimodal reasoning preference data and high annotation cost. Existing multimodal preference datasets primarily address hallucination issues and focus on natural images and perception data, lacking scientific images and reasoning data. Annotating these types of data requires human annotators to carefully compare the given reasoning processes, making it both time-consuming and costly. (2) Lack of open-source methods for improving multimodal reasoning via PO. Although previous works have explored fine-tuning MLLMs using feedback from various sources, these models typically exhibit performance gains on hallucination benchmarks, with little enhancement in general reasoning abilities. Thus, leveraging PO to improve multimodal reasoning capabilities remains largely under-explored.

This work addresses these challenges from both the data and model sides. (1) On the data side, we design an automated preference data construction pipeline to create MMPR, a high-quality, large-scale multimodal reasoning preference dataset. (2) On the model side, we explore various PO methods with MLLMs, introducing a simple yet effective method, termed Mixed Preference Optimization (MPO), which boosts multimodal CoT performance without the requirement for a reward model.

Specifically, we propose a continuation-based pipeline called Dropout Next Token Prediction (DropoutNTP) for samples lacking clear ground truth and a correctness-based pipeline for samples with clear ground truth. In DropoutNTP, the responses generated by InternVL2-8B are considered as positive samples. For a given chosen response, we truncate it by half and then prompt InternVL2-8B to complete the remaining portion of the truncated answer without access to the image input. This generated completion serves as the rejected answer for the paired sample. Experimental results in Section 5.2 demonstrate that this straightforward method achieves comparable performance in reducing hallucinations compared to the divide-and-conquer method proposed in RLAIF-V . In the correctness-based pipeline, multiple solutions to each question are sampled from InternVL2-8B. Solutions matching the ground truth answer are used as chosen responses, while those that do not are used as rejected responses.

Additionally, we propose the MPO method. The key insight behind this algorithm is that an effective PO process should enable the model to learn the relative preference between pairs of responses, the absolute quality of individual responses, and the process for generating preferred responses. Compared to previous multimodal PO methods , our approach excels in the following aspects: (1) Efficient automated data construction pipeline: Our pipeline enables high-quality preference pair generation at a controlled cost. (2) Effectiveness across diverse domains: Models fine-tuned with our data and approach show superior performance across reasoning, question-answering, and hallucination benchmarks. (3) Improvements over SoTA settings: Our results are based on InternVL2-8B, one of the leading open-source MLLMs, further highlighting the potential of our method.

In summary, our main contributions are as follows:

(1) We propose an efficient preference data construction pipeline. Based on this pipeline, we create MMPR, a high-quality, large-scale multimodal reasoning preference dataset containing approximately 3 million samples.

(2) We introduce MPO, an effective PO algorithm designed to improve the reasoning abilities of MLLMs. The resulting model, InternVL2-8B-MPO, exhibits enhanced multimodal reasoning ability and fewer hallucinations compared to its baseline model (i.e., InternVL2-8B).

(3) We conduct extensive experiments to explore practical approaches for improving multimodal reasoning via PO. Results show that PO significantly improves reasoning abilities over SFT. Notably, the proposed InternVL2-8B-MPO achieves an accuracy of 67.0 on MathVista , outperforming InternVL2-8B by 8.7 points and achieving performance comparable to the 10×\times larger InternVL2-76B.

Related Work

Multimodal Large Language Models. With advancements in LLMs, significant progress has also been made in MLLMs. To leverage the abilities of pre-trained LLMs and Vision Foundation Models (VFMs) , a series of works employ a connector to align their latent space, achieving promising performance at a controllable cost. Besides, another series of works extend pre-trained LLMs with additional fusion layers for vision features, reducing the number of visual tokens required by LLMs while introducing extra training costs. Recently, there have been explorations into vision encoder-free architectures , which consists of a single transformer model that jointly processes visual and textual information without a separate encoder. In addition to exploring model architectures, recent works also try to construct high-quality training data to improve multimodal reasoning abilities. Despite these advancements, MLLMs typically rely on a training paradigm comprising pre-training and supervised fine-tuning, which suffers from the curve of distribution shift and exhibits limited multimodal reasoning abilities. In this work, we conduct a systematic study on using preference optimization to enhance the multimodal reasoning ability of MLLMs.

Preference Optimization. Preference optimization (PO) is a crucial technique for advancing LLMs and MLLMs. Specifically, Reinforcement Learning from Human Feedback (RLHF) uses human preferences as a reward signal to fine-tune models, aligning them with human preferences. InstructGPT employs a reward model as a proxy for human preferences and maximizes this reward via the PPO algorithm , improving the model’s ability to follow user intent and become more helpful, honest, and harmless (3H). PPO-Max carefully explores the implementation details of PPO, proposing a more stable version of the algorithm. Additionally, DPO proposes an efficient PO algorithm based on the Bradley-Terry model , removing the need for an explicit reward model. Subsequent works have further analyzed and refined this method from various perspectives. In natural language processing, a series of works have explored how to leverage PO to enhance reasoning ability. In the multimodal field, however, most methods primarily focus on reducing hallucination, leaving the potential for PO to improve multimodal reasoning ability under-explored. This work demonstrates that PO not only mitigates hallucinations but also strengthens multimodal reasoning abilities, highlighting its broader applicability in MLLM development.

Scalable Multimodal Preference Dataset Generation

To address the scarcity of multimodal preference data, we introduce a scalable data construction pipeline. Based on this pipeline, we construct a million-level MultiModal PReference dataset (MMPR).

Definition. Each data sample in our MMPR consists of an image I∈II\in\mathcal{I}, an instruction x∈Xx\in\mathcal{X}, a chosen response yc∈Ypy_{c}\in\mathcal{Y}_{p}, and a rejected response yr∈Yny_{r}\in\mathcal{Y}_{n}, where ycy_{c} is preferable to yry_{r}. The image sets I\mathcal{I} and instruction sets X\mathcal{X} are collected from existing datasets. Yp\mathcal{Y}_{p} and Yn\mathcal{Y}_{n} represent the positive and negative response set, respectively. Given a certain image II and instruction xx, we sample the candidate response yy from an initial instruction model M0M_{0} as follows:

where M0(y∣x,I)M_{0}(y\mid x,I) represents the response distribution of M0M_{0} conditioned on image II and instruction xx.

For instructions with clear ground truths, the model is prompted to first provide the reasoning process and then give the final answer in the format like “Final Answer: ***”. Responses matching the ground truth answer constitute the positive set Yp\mathcal{Y}_{p}, while those that do not match make up the negative set Yn\mathcal{Y}_{n}. Additionally, responses that fail to provide a clear final answer are also merged into Yn\mathcal{Y}_{n}. Given these responses labeled as positive or negative, we build the preference pairs by selecting a chosen response ycy_{c} from Yp\mathcal{Y}_{p} and a negative response yry_{r} from Yn\mathcal{Y}_{n}.

For instructions without clear ground truths, we propose a simple yet effective method: Dropout Next-Token Prediction (Dropout NTP). Specifically, we directly consider all responses generated from equation 1 as positive set Yp\mathcal{Y}_{p}. To generate the negative set Yn\mathcal{Y}_{n}, we sample a response yy from Yp\mathcal{Y}_{p} and drop the last half of this response. The model is required to complete the remained response as follows:

Compared with previous methods, our data engine is as effective as the more complex divide-and-conquer method proposed in RLAIF-V (see the experimental results in Section 5.2.2), while more efficient. Taking data generation for M3CoT as an example, our pipeline incurs a token cost of 571.2 per preference pair, compared to 992.7 tokens for the divide-and-conquer approach used in RLAIF-V. Thus, the cost of our pipeline is only 57.5% of that of RLAIF-V.

2 Multimodal Preference Dataset

Dataset Statistics. Using this pipeline, we build a large-scale multimodal preference dataset, MMPR. Data examples are presented in Figure 2. See more examples in the Appendix. This dataset comprises approximately 750K samples without clear ground truths and 2.5M samples with clear ground truths. For samples without clear ground truths, each instruction averages 25.0 tokens, while the chosen and rejected responses average 211.4 and 171.2 tokens, respectively. The longest chosen and rejected responses consist of 1,342 and 1,642 tokens, respectively, whereas the shortest chosen and rejected responses contain 20 and 17 tokens, respectively. For samples with clear ground truths, the average instruction length is 79.5 tokens, with the chosen and rejected responses averaging 300.0 and 350.5 tokens, respectively. The longest chosen and rejected responses are composed of 2,018 and 4,097 tokens, while the shortest responses contain 32 and 33 tokens, respectively.

Data Source. As shown in Table 1, to ensure the diversity of instructions and images, we collect samples from diverse domains, including general visual question answering (VQA) , science , chart , mathematics , OCR , and document . Notably, when constructing open-ended samples, we collect instructions from all the data sources mentioned above and prompt the model to answer the original question without additional requirements. On the other side, when building samples through the correctness-based pipeline, we exclude questions from general VQA and document sources, as verifying the correctness of the generated answers using heuristic rules is challenging for datasets in these domains. For example, the ground truths in VQAv2 consist of a single word or phrase, which may lead to false-negative responses when the model outputs a complete sentence or a synonym as the final answer. Such false-negative responses can negatively impact training effectiveness.

Improved Multimodal Large Language Model with Preference Optimization

To enhance the multimodal reasoning capabilities of MLLMs, we propose mixed preference optimization (MPO), a method that blends supervised fine-tuning (SFT) loss with various preference optimization losses to enhance training effectiveness. Additionally, we investigate different Chain-of-Thought (CoT) approaches with multimodal input to improve reasoning performance.

We observed that when MLLMs are trained on large-scale preference datasets using direct preference optimization (DPO), they might fail to generate reasonable rationales and produce gibberish. This phenomenon aligns with the analysis presented in Smaug . To address this issue, we introduce the MPO in this work, aiming to learn the relative preference between pairs of responses, the absolute quality of individual responses, and the process for generating preferred responses.

Training Objective. MPO is defined as a combination of preference loss Lp\mathcal{L}_{p}, quality loss Lq\mathcal{L}_{q}, and generation loss Lg\mathcal{L}_{g}, which can be formulated as follows:

where w∗w_{*} represents the weight assigned to each loss component. In this work, we empirically compare different variants of preference loss . Based on the experimental results, we use DPO as our preference loss and BCO as our quality loss.

Preference Loss. The DPO serves as the preference loss to enable the model to learn the relative preference between chosen and rejected responses. DPO eliminates the requirement of training an explicit reward model based on the assumption of the Bradley-Terry model and optimizes the following loss function:

where β\beta is the KL penalty coefficient, and xx, ycy_{c}, and yry_{r} are user query, chosen response, and rejected response, respectively. The policy model πθ\pi_{\theta} is initialized from model π0\pi_{0}.

Quality Loss. The BCO loss is employed as the quality loss, which helps the model to understand the absolute quality of individual responses. This algorithm trains a binary classifier, where the logit serves as a reward and effectively maps the chosen response to 1 and the rejected response to 0. The loss function is defined as:

where Lq+\mathcal{L}_{q}^{+} and Lq−\mathcal{L}_{q}^{-} represent the loss for chosen and rejected responses, respectively. They are calculated independently, requiring the model to differentiate the absolute quality of individual responses. The loss terms are given by:

where δ\delta represents the reward shift, calculated as the moving average of previous rewards to stabilize training.

Generation Loss. The SFT loss is used as the generation loss to help the model learn the generation process of preferred responses. The loss function is defined as:

2 Chain-of-Thought with Multimodal Input

During the data sampling process, we require the model to provide a detailed CoT reasoning process instead of directly answering the final answer. For most samples, we sample the responses using the prompt shown in the bottom case of Figure 2, which requires the model to perform a step-by-step analysis. Considering that multimodal models involve non-textual inputs, we further introduce the following CoT methods: (1) Background Knowledge-based CoT: The model first introduces relevant background knowledge related to the problem or image, followed by reasoning steps and the final answer. This approach is applied to samples from the science domain. (2) Visual Content-based CoT: The model begins by analyzing the visual contents in the image, then proceeds with reasoning and the final answer. This method is used for samples from chart, OCR, and document domains. (3) Grounded CoT: The model generates a text response while simultaneously linking all referenced objects in the response to corresponding regions in the image. This approach is applied to general VQA domain samples. Responses generated by these above CoT methods are mixed with those sampled using the prompt shown in the bottom case of Figure 2. These approaches not only effectively integrate multimodal information into the reasoning process but also enhance data diversity. Furthermore, including the background knowledge and visual contents at the start of responses also improves the quality of the negative responses generated by DropoutNTP, preventing a significant quality gap between positive and negative samples that could reduce training effectiveness.

Experiments

In this section, we compare our InternVL2-8B-MPO with leading MLLMs on multimodal reasoning , complex Visual Question Answering (VQA) , and hallucination evaluation tasks.

Benchmarks. For the multimodal reasoning task, we evaluate our model on three benchmarks: (1) M3CoT , a comprehensive benchmark designed to evaluate the multimodal CoT reasoning abilities of models. (2) MathVista , a widely-used benchmark for evaluating multimodal mathematical reasoning capabilities. (3) MathVision , which collects evaluation data from real math competitions and presents a greater challenge compared to MathVista. We report accuracy for these benchmarks.

For the complex VQA task, we evaluate our model on two benchmarks: (1) MM-Vet , which evaluates the model’s ability to engage in visual conversations across a diverse range of tasks. (2) LLaVA-Bench , a commonly-used benchmark for assessing multimodal conversation, detailed description, and complex reasoning capabilities with open-ended questions. Both benchmarks use GPT-4 to evaluate the correctness and helpfulness of responses. We report the overall score for these benchmarks.

For the hallucination evaluation task, we evaluate our model on three benchmarks: (1) POPE , which measures the hallucination level of object existence using Yes/No questions. We report the F1 score for this benchmark. (2) CRPE , which measures the hallucination level of the relation between objects using multiple-choice questions. We report accuracy for this benchmark. (3) MMHal-Bench , which consists of open-ended questions where GPT-4 compares model outputs to human responses, assessing hallucination rate and informativeness. We report the overall score for this benchmark.

Results. As shown in Table 2, our InternVL2-8B-MPO achieves superior performance across all benchmarks, particularly excelling in multimodal reasoning tasks. On the MathVista benchmark, our model achieves an accuracy of 67.0%, outperforming InternVL2-8B by 8.7 points and achieving performance comparable to the 10×\times larger InternVL2-76B. On the MathVision benchmark, our model achieves an accuracy of 25.7%, establishing a new state-of-the-art performance among open-source models. These results demonstrate the effectiveness of our preference optimization approach in enhancing multimodal reasoning capabilities. Additionally, on the POPE benchmark, our model exhibits a 1.2-point improvement over InterVL2-8B, demonstrating the effectiveness of the perception data contained in our MMPR dataset to mitigate hallucinations. Furthermore, our model also shows superior performance compared to the InternVL2-8B on complex VQA benchmarks, indicating that the general abilities of our model are also improved, benefiting from enhanced reasoning abilities and mitigated hallucinations.

2 Ablation Study

In this section, we present ablation studies to analyze the effects of preference optimization and SFT on multimodal reasoning abilities. Additionally, we compare our proposed DropoutNTP method with the divide-and-conquer approach from RLAIF-V , demonstrating the effectiveness of our approach. Furthermore, we conduct extensive experiments to analyze the effects of different preference optimization algorithms. We also present analysis of the effects on text-only performance.

To compare the impact of MPO and SFT on improving multimodal reasoning ability, we use the chosen responses in MMPR as SFT data to fine-tune InternVL2-8B. As shown in Table 3, the results indicate that the model trained with MPO consistently outperforms that trained with SFT across all benchmarks. For example, the MPO-trained model achieves a score of 79.2 on the multimodal reasoning benchmark M3CoT, surpassing its SFT counterpart by 11.4 points. Furthermore, the MPO-trained model also performs better on the general benchmark (MMVet) and the hallucination benchmark (POPE). Notably, the SFT-trained model performs worse with CoT responses than with direct-answer responses on MMVet and POPE, demonstrating that SFT alone is insufficient to enhance multimodal CoT abilities. These results demonstrate that while SFT provides moderate improvement, preference optimization is more effective in improving the overall performance of the model.

2.2 Comparison with RLAIF-V

Here, we compare our proposed Dropout Next-Token Prediction (Dropout NTP) method with the divide-and-conquer approach from RLAIF-V . To ensure a fair comparison, we use the same prompts and chosen responses as in RLAIF-V and replace the rejected responses with those generated by continuation without image input. Following RLAIF-V, we report the hallucination rates in response-level (Resp.) and mention-level (Ment.) for Object HalBench and overall score and hallucination rates (Hall.) for MMHal-Bench . As shown in Table 4, the model trained with our data achieves performance comparable to that of the model trained with RLAIF-V, demonstrating the effectiveness of our method. Specifically, the response-level hallucination rate of the model trained with our data on Object HalBench is 7.6, compared to 7.3 for its counterpart. Besides, this model achieves a score of 3.6 on the MMHal-Bench, compared to 3.5 for its counterpart. Note that our method requires the model to generate only a single continuation for each sample, while RLAIF-V requires the model to decompose the response into atomic claims and then verify each one individually. Therefore, our method is more efficient. A quantitative analysis is provided in Section 3.1.

2.3 Effects of optimization algorithms

Here, we empirically compare the effectiveness of different optimization algorithms, including (1) DPO , which directly fine-tunes the model on an offline preference dataset without explicitly constructing a reward function. (2) RSO , which applies a hinge loss on the normalized likelihood instead of the sigmoid loss used in DPO. (3) IPO , which introduces a modified loss function to address overfitting in DPO by averaging log-likelihoods and controlling the gap between chosen and rejected completions via a beta parameter. (4) cDPO , which is a modification of the DPO loss that accounts for potential label noise in preference data. (5) RobustDPO , which provides an unbiased estimate of the DPO loss designed to handle preference noise in data. Similar to cDPO, it assumes that labels are noisy with a certain probability. (6) BCO , which introduces a binary classifier trained to output logits used as reward values. (7) SPPO , which iteratively pushes chosen rewards toward 1/2 and rejected rewards toward -1/2 to approximate a Nash equilibrium, aiming to reduce data sparsity issues. (8) AOT , which applies Distributional Preference Alignment via Optimal Transport. (9) TR-DPO , which adds synchronization between the model and a reference model every few steps to mitigate overfitting during DPO training. (10) ORPO , a reference model-free preference optimization algorithm that uses a log odds ratio penalty appended to the NLL loss, allowing for preference-aligned fine-tuning without an additional preference alignment phase. For all algorithms, we set the learning rate to 5e-65e\text{-}6 and use the hyper-parameters suggested in their corresponding paper. Additionally, we extend these algorithms with SFT loss to analyze its impact. The SFT model trained with the chosen responses of the reasoning preference data is also included as a baseline.

Notably, most current benchmarks lack corresponding in-distribution training samples, and the data distribution of our MMPR may differ from that of these benchmarks. This discrepancy can introduce additional variability when analyzing the impact of different optimization algorithms on training results. Therefore, we use the training and validation sets of M3CoT for ablation studies.

The visualization results are illustrated in Figure 3 and the numerical results are presented in Table 6 and 7. We can observe that almost all preference optimization methods outperform their SFT counterpart in both the Direct and CoT settings. However, DPO and its variants struggle to enhance the CoT reasoning abilities of the model as the resulting models exhibit trivial or no improvement when answering with CoT reasoning responses compared to direct-answer responses. On the other hand, when combining SFT Loss with these DPO variants, all algorithms are able to improve the model’s CoT reasoning abilities, demonstrating that the SFT loss is a key component for enhancing CoT reasoning abilities. Additionally, models trained with TR-DPO, a DPO variant that updates the reference model every few steps, perform much worse when using CoT reasoning compared to direct-answer responses. Similarly, the model trained with ODPO, a reference-model-free method, achieves worse overall performance compared to other methods extended with SFT Loss. These results indicate that the reference model constraint on policy updates is crucial for enhancing overall reasoning abilities, and the reference model should remain frozen during training. Notably, models trained with DPO+ and BCO+ exhibit the best CoT performance among existing algorithms. Therefore, we use DPO and BCO as the preference loss and quality loss. The resulting algorithm (i.e., MPO) further improves the overall performance.

3 Effects on text-only performance

We evaluate the text-only performance of our models on a series of benchmarks and report the average performance across them. As shown in Table 5, although our MMPR dataset does not include any text-only data, the MPO-trained model achieves superior average performance on these benchmarks compared to the baseline model. The most significant improvements are observed on TheoremQA and IFEval. Specifically, our model trained with MPO achieves an accuracy of 20.8 on TheoremQA, a benchmark consisting of complex science problems, outperforming the baseline model by 5.2 points and the SFT counterpart by 5.0 points. Additionally, since our dataset considers responses that fail to follow instructions as negative samples when constructing data using our correctness-based pipeline, our model also exhibits enhanced instruction-following abilities on IFEval, outperforming the baseline model by 4.1 points and the SFT counterpart by 2.8 points.

Conclusion

In this work, we introduce a preference optimization (PO) process to enhance the multimodal reasoning capabilities of MLLMs. On the data side, we design an automated pipeline for preference data construction, which is applicable to instructions both with and without clear ground truths. Using this pipeline, we create MMPR, a high-quality, large-scale multimodal reasoning preference dataset. On the model side, we propose a simple yet effective method called Mixed Preference Optimization (MPO). This algorithm aims to learn the relative preference between pairs of responses, the absolute quality of individual responses, and the process for generating preferred responses. The resulting model, InternVL2-8B-MPO, exhibits enhanced multimodal reasoning ability and fewer hallucinations compared to its baseline model (i.e., InternVL2-8B). We hope this study could inspire further advancements in MLLMs.

References

Implementation Details

During the construction of samples with clear ground truths, we sample at most 3232 reasoning processes and construct at most 1515 preference pairs for each query. When constructing data using DropoutNTP, we truncate the original response by half and ask InternVL2-8B to complete the response without the image input. Our ablation studies in Section 8.2 show that truncating the original response by 25% or 75% has negative effects on the final performance. We set the temperature to 1.01.0 during sampling to ensure response diversity. Besides, the maximum tiles for dynamic resolution are set to 66 for the general VQA domain and 1212 for OCR-, document-, and chart-related domains.

During the MPO process, the global batch size is set to 256256 during training. We employ the AdamW optimizer with the β1\beta_{1} of 0.90.9, the β2\beta_{2} of 0.9990.999, and the weight decay of 0.050.05. The learning rate is initialized as 5e-65e\text{-}6. The training phases include a linear warmup that lasts until the first 5% of training steps. The warmup is followed by a cosine decay strategy with a minimum learning rate of 0. The KL penalty coefficient β\beta is set to 0.10.1. For the Equation 3, we set wpw_{p} to 0.80.8, wqw_{q} to 0.20.2, and wgw_{g} to 11. The model is initialized from InternVL2-8B , and all parameters are trainable during training. We train the model for 1 epoch.

More Ablation Studies

In this section, we present the numerical experimental results of ablation studies on the effects of different preference optimization algorithms in Table 6 and Table 7. We define Δ\Delta as the performance gap between CoT reasoning responses and direct-answer responses to quantitatively assess the effects of different preference optimisation algorithms on CoT reasoning abilities. Our results indicate that introducing an additional SFT loss can significantly improve the CoT performance compared to each algorithm’s vanilla counterpart. Note that, to reduce computational costs, we only extend the DPO variants, which exhibit superior performance in Table 6 compared to DPO, with SFT Loss.

In addition to the ablation studies based on M3CoT, we also present the performance of models trained with DPO+ and BCO+ using our MMPR, as shown in Table 8. The experimental results show that models trained with MPO exhibits superior overall performance compared to those trained with DPO+ and BCO+.

2 Ablation Studies on DropoutNTP

Here, we present the ablation results for the Dropout Ratio (DR) in our proposed DropoutNTP. By default, we set DR to 0.50.5, which means that we truncate the positive response by half. Notably, setting DR to 0.250.25 means using the first quarter of the positive responses for continuation. Following the experimental settings in Section 5.2.2, we replace the negative samples in RLAIF-V with the completions based on different dropout ratios. As shown in Table 9, the model trained with data generated using a DR of 0.750.75 performs the worst. We attribute this to the fact that, with the first three-quarters of the prefix being identical, the difference in quality between the chosen and rejected responses becomes less apparent, reducing training effectiveness. Additionally, the model trained with a DR of 0.250.25 performs worse than that trained with a dropout ratio of 0.50.5. We believe this is because the majority of the content in the rejected responses is generated without image input, resulting in noticeably lower quality compared to the chosen responses, which similarly hampers the training effectiveness. Therefore, we set the DR to 0.50.5.

3 Effects of data scale.

To evaluate the effects of the data scale, we train the model with different amounts of preference reasoning data sampled from M3CoT . The M3CoT training set contains 7,861 samples annotated with corresponding rationales. To control the data volume, we adjust the maximum number of preference pairs generated for each sample, resulting in datasets of different sizes: 10K, 40K, 70K, and 100K. As illustrated in Figure 4(a), model accuracy consistently improves with the increasing data volume. As the data volume rises to 100K, the model achieves its highest accuracy of 76.4 when directly answering the final answer and 78.9 when answering with CoT. Furthermore, both the Direct and CoT performance exhibit a positive correlation between data scale and accuracy, with the CoT performance achieving higher performance across all scales. These results highlight the importance of scaling up reasoning preference data to improve model performance.

4 Effects of hyper-parameters.

We conduct ablation studies on M3CoT to study the impact of the hyper-parameters, including learning rate, PO coefficient wp,wqw_{p},w_{q}, and SFT coefficient wgw_{g}. For the PO coefficient, we control the sum of wpw_{p} and wqw_{q} to equal 1.0 and adjust different proportions. Unless specifically mentioned, we set the learning rate to 5e-65e\text{-}6, wpw_{p} to 0.80.8, wqw_{q} to 0.20.2, and wgw_{g} to 11. As shown in Figure 4(b), the learning rate significantly affects the model’s performance. With a relatively low learning rate of 5e-75e\text{-}7, the model shows moderate improvement. As the learning rate increases to 5e-65e\text{-}6, the model’s performance improves further, reaching optimal results across the tested learning rates and surpassing the baseline by 19.6 points. However, further increasing the learning rate to 5e-55e\text{-}5 causes a drastic performance drop, suggesting that a higher learning rate may lead to overfitting or instability in training. Additionally, the PO coefficient w0,w1w_{0},w_{1} and SFT coefficient w2w_{2} are crucial. As shown in Figure 4(c) and 4(d), the model achieves optimal performance with wpw_{p} set to 0.80.8, wqw_{q} set to 0.20.2, and wgw_{g} set to 11. Notably, when wgw_{g} is set to 0.010.01, the performance of the CoT approach is inferior to that of directly answering the final answer, indicating the importance of the SFT Loss during the direct preference optimization.

More Data Examples in MMPR

In this section, we provide data examples in MMPR for each task described in Table 1. Specifically, Figure 5(a) to 5(f) are examples from data constructed using DropoutNTP, while Figure 5(g) to 5(j) are examples from data constructed using correctness-based pipeline.