OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning

Zhaochen Su, Linjie Li, Mingyang Song, Yunzhuo Hao, Zhengyuan Yang, Jun Zhang, Guanjie Chen, Jiawei Gu, Juntao Li, Xiaoye Qu, Yu Cheng

Introduction

“The eye sees only what the mind is prepared to comprehend.”

Recent advances in large vision-language models (LVLMs) have significantly expanded the capabilities of AI agents to jointly reason over visual and textual inputs Liu et al. (2023a); Zhu et al. (2023); Su et al. (2024a). By leveraging techniques such as chain-of-thought (CoT) prompting (Wei et al., 2022; Su et al., 2024b), these models have achieved impressive performance on a broad range of multimodal tasks, such as visual question answering (Antol et al., 2015), mathematical reasoning (Lu et al., 2024), and image captioning (Sharma et al., 2018). However, most current approaches still rely primarily on textual intermediate reasoning, even when dealing with inherently visual problems.

In contrast, human reasoning is often deeply intertwined with visual cognition (Zhang and Norman, 1994; Larkin and Simon, 1987). People not only describe what we see but also think with images, using sketches (Goel, 1995), highlights (Kosslyn, 1994), and spatial cues (Tversky, 2005; Tversky et al., 2002) to externalize, decompose, and manipulate complex visual information. For example, when solving a geometry problem, people will draw auxiliary lines or mark key points on a diagram to reveal hidden relationships and guide reasoning. These biological mechanisms motivate employing visual tools as cognitive scaffolds for LVLMs. By integrating tools directly into the reasoning loop, we enable models to iteratively manipulate and interpret visual content, providing a more grounded and interpretable decision-making pathway. This paradigm shift from pure text-based reasoning to tool-augmented visual cognition holds promise for solving tasks that demand fine-grained spatial understanding, iterative perception, and precise interaction with visual content.

Recent efforts have begun to explore tool-augmented multimodal reasoning (Hu et al., 2024; Ma et al., 2024a; Fu et al., 2025) by equipping agents with the ability to interact with external visual tools, compose intermediate visual representations, and learn action trajectories through synthetic supervision (Liu et al., 2023b; Hu et al., 2024; Ma et al., 2024b). While demonstrating tool integration potential, these SFT-centric approaches, typically relying on orchestrated tool-use sequences from static datasets, limit holistic learning across the tool-use lifecycle. These approaches gives rise to several fundamental challenges: ❶ Heterogeneous tool definitions and interfaces: Tools with identical names (e.g., “segment” or “grounding”) often differ in behavior due to backend implementations or task-specific assumptions, hindering standardization and reproducibility. ❷ High cost of trajectory generation: Producing training data for tool-based reasoning is resource-intensive, often relying on manual templates or brittle heuristics that limit scalability and accuracy verification. ❸ Limited training generalization: Existing methods typically adopt SFT on offline trajectories. However, SFT alone struggles to generalize to unseen tools or tasks, and lacks mechanisms for exploration and dynamic adaptation (Su et al., 2024c; Jin et al., 2025; Chu et al., 2025a; DeepSeek-AI, 2025).

To address these challenges, in this paper, we introduce OpenThinkIMG, the first comprehensive end-to-end framework unifying these critical stages for tool-augmented LVLMs. Specifically, OpenThinkIMG provides a unified infrastructure for standardizing heterogeneous tool interfaces, scaling the generation of tool-use trajectories, and supporting efficient training of multimodal agents. Beyond traditional SFT approaches, we further propose V-ToolRL, a reinforcement learning framework that enables models to autonomously explore and discover optimal tool usage strategies with vision tools. By tightly integrating flexible tool management, scalable trajectory synthesis, and dynamic agent adaptation, OpenThinkIMG offers a practical foundation for building next-generation LVLMs with enhanced visual reasoning capabilities. Our main contributions are summarized as follows:

We introduce OpenThinkIMG, the first open and extensible end-to-end framework for tool-augmented LVLMs. It features a unified registry for diverse vision tools and backbone models, a distributed deployment strategy for efficient and scalable tool inference, and an integrated E2E training pipeline that incorporates our proposed novel V-ToolRL methodology for adaptive tool use. All code and resources are publicly available and will be actively maintained to foster community collaboration and further development in tool-augmented reasoning.

We propose a scalable and adaptable three-stage pipeline for constructing high-quality vision tool-use trajectories. This pipeline leverages the model’s capabilities for initial action planning, performs automated tool call completion and rationale parsing, and incorporates multi-stage filtering with rule-based validation and human oversight to ensure data quality for both supervised fine-tuning and reinforcement learning.

We empirically validate V-ToolRL on complex chart reasoning tasks. Our approach boosts the performance of a 2B parameter base model by +29.83 accuracy points and surpasses larger 8B/13B open-source tool-augmented agents by an average of 12.7 points. Detailed experiments and qualitative studies further illustrate the learned tool-use efficiency, the development of complex reasoning narratives, and the superior interpretability of our method.

OpenThinkIMG Framework

In this section, we will detail the architecture of OpenThinkIMG, a comprehensive, community-driven framework designed to streamline the integration of vision tools, scale the synthesis of tool-use trajectories, and support efficient training of multimodal agents. It encompasses a unified registry for tools and models, a distributed deployment strategy for dynamic inference, and an integrated training pipeline featuring both supervised fine-tuning and our proposed V-ToolRL for learning adaptive tool invocation. The overall architecture and process flow are illustrated in Figure 1.

Effectively tackling diverse visual reasoning tasks necessitates a versatile suite of tools. To address this challenge, OpenThinkIMG provides a unified registry for the seamless integration of vision tools and backbone models, requiring minimal boilerplate. The framework, therefore, incorporates a curated selection of vision tools designed to address specific facets of visual interaction and reasoning. Table 1 offers a comprehensive summary of each tool’s detailed parameters and specifications, while their core functionalities and typical use cases are detailed below:

GroundingDINO (Liu et al., 2024a): This tool bridges language and visual perception by performing text-driven object detection. It takes an input image IinI_{in} and a textual query qtextq_{text} to locate instances of described objects, outputting their bounding boxes Bout\mathcal{B}_{out}. It is indispensable for tasks requiring the model to answer “Where is X?” or “Find all Y” based on visual content.

SAM (Segment Anything Model) (Kirillov et al., 2023): Motivated by the need for precise, object-agnostic segmentation, SAM generates fine-grained segmentation masks moutm_{out}. It typically takes an input image IinI_{in} and a prompt like an input bounding box binb_{in} (or points). This is crucial for isolating specific objects for detailed analysis or manipulation, regardless of object class, especially when precise boundaries are needed.

OCR (Optical Character Recognition): Designed to extract and understand textual information embedded within images, OCR processes an input image IinI_{in} to identify and transcribe text. It outputs the extracted text toutt_{out} along with the bounding boxes Bout\mathcal{B}_{out} of the text regions. This is essential for tasks involving reading labels on charts, signs, documents, or any scenario where textual content in an image is relevant.

Crop: This tool allows focusing processing or attention on a specific sub-region of an image. Given an input image IinI_{in} and the bounding boxes BinB_{in}, it extracts a rectangular sub-region, outputting the cropped image IcropI_{crop}. It is useful for isolating a region of interest for subsequent, more detailed analysis by other tools or when only part of the image is relevant.

Point: Intended for precisely identifying a single location or object based on descriptive language, the Point tool takes an input image IinI_{in} and a textual description qtextq_{text}. It localizes the specified object or point of interest and returns its coordinates poutp_{out}. This is valuable for tasks that require pinpointing a specific item, such as “mark the highest peak”.

DrawHorizontalLineByY / DrawVerticalLineByX: These tools visually aid reasoning by adding reference markers to an image. They take an input image IinI_{in} and a Y or X coordinate value cinc_{in}, respectively, and output an annotated image IannotI_{annot} with the corresponding horizontal or vertical line drawn. They are particularly useful in chart and graph analysis for marking thresholds or comparing values.

ZoomInSubplot: To enable detailed examination of specific parts within complex visuals, this tool creates magnified views (subplots) Isubplot\mathcal{I}_{subplot}. It takes an input image IinI_{in} and a textual description qtextq_{text} (or coordinates cinc_{in}) identifying the region to zoom into. It is beneficial when analyzing images with multiple distinct areas requiring closer inspection.

SegmentRegionAroundPoint: This tool is used to refine segmentation locally or isolate a small feature with high precision. Starting from an input image IinI_{in} and a designated point coordinate pinp_{in}, it generates or refines a segmentation mask mlocalm_{local} specifically around that point. This is useful for obtaining precise masks for small objects or refining coarse segmentations.

To streamline model loading, we employ the Transformers library (Wolf et al., 2020) to load pre-trained models and initialize parameters. Closed-source models are loaded from the OpenAI repository. At present, we support the Gemini, ChatGPT, Qwen-2VL, and Qwen-2.5VL series models. Furthermore, OpenThinkIMG includes streamlined deployment modules for both vision tools and models, and we will continue to expand its repertoire of supported components in the future.

2 Vision Tool Deployment and Inference

A key architectural choice in OpenThinkIMG is the distributed deployment of vision tools, contrasting with prior approaches that often load all tools into a single memory space (Wu et al., 2023a; Ma et al., 2024b). This modular design enhances scalability, fault isolation, and allows for independent updates and resource allocation for each tool. Specifically, each vision tool Tk∈TsuiteT_{k}\in\mathcal{T}_{suite} (where Tsuite\mathcal{T}_{suite} is the suite of available tools), is deployed as an independent, containerized service SkS_{k}, listening on a dedicated local network port.

To effectively manage these distributed services, a Tool Controller is designed, which orchestrates the entire tool invocation lifecycle. While the controller handles service registration and health monitoring, its core function is dynamic inference-time orchestration. During inference, upon an LVLM identifying a need for tool assistance based on the current input, such as a question QQ and image II, it formulates a planned action ata_{t}. This plan typically specifies the tool to be called TkT_{k} and its arguments, which are derived from the LVLM’s internal reasoning state RLVLMR_{LVLM} and the input (Q,I)(Q,I). The Tool Controller receives this planned action ata_{t}. It then parses the request, determines an efficient execution strategy (potentially parallelizing if ata_{t} represents multiple independent tool calls), and subsequently dispatches ata_{t} to the corresponding service SkS_{k}. The service executes the tool, effectively performing a step in the tool rollout process O(⋅∣at,(Q,I))O(\cdot\mid a_{t},(Q,I)), yielding an output ot←Sk(at)o_{t}\leftarrow S_{k}(a_{t}). If multiple tools are called, their outputs are aggregated by the controller into a set of outcomes ωt=(ot,1,ot,2,… )\omega_{t}=(o_{t,1},o_{t,2},\dots). Finally, the controller augments the LVLM’s current reasoning context (e.g., RLVLMR_{LVLM}) with ωt\omega_{t} (or oto_{t} if a single tool) to form an updated context Caugt=(RLVLM,ωt)C_{aug_{t}}=(R_{LVLM},\omega_{t}). This CaugtC_{aug_{t}} is returned to the LVLM for subsequent reasoning steps or final response generation, enabling an iterative, multi-step problem-solving process.

3 V-ToolRL: Reinforcement Learning with Vision Tools

The OpenThinkIMG architecture detailed above provides the robust infrastructure for flexible tool deployment and dynamic inference. However, to empower the LVLM to learn how and when to strategically leverage this toolset for optimal task completion, a dedicated learning paradigm is essential. In this section, we will introduce our proposed novel methodology called V-ToolRL, which consists of two modules: a cold-start module for initializing vision tool invocation and a reinforcement learning module for adaptive tool usage.

To bootstrap basic vision tool invocation, we first perform supervised fine-tuning on the batch-generated trajectories. Each trajectory is defined as:

where at(i)a_{t}^{(i)} denotes the planned action at step tt and ot(i)o_{t}^{(i)} the corresponding tool output for the ii-th example. Based on the trajectory generation procedure described in Section 3, we construct the training dataset \mathcal{D}=\bigl{\{}(Q^{(i)},I^{(i)},\tau^{(i)})\bigr{\}}_{i=1}^{N}, where Q(i)Q^{(i)} is the ii-th question prompt, I(i)I^{(i)} the associated input image, τ(i)\tau^{(i)} the action–output trajectory of length n(i)n^{(i)}, and NN the total number of examples. During the Cold-Start stage, the model learns to generate the full trajectory τ(i)\tau^{(i)} conditioned on (Q(i),I(i))(Q^{(i)},I^{(i)}). We optimize the cross-entropy loss:

3.2 Reinforcement Learning for Adaptive Tool Usage

We train V-ToolRL using the Group-wise Proximal Policy Optimization (GRPO) algorithm (Shao et al., 2024), extended to account for vision-tool rollouts. Concretely, for each question q∼P(Q)q\sim P(Q) we sample a group of GG candidate action trajectories:

and then execute each planned action sequence via our vision tools to obtain the corresponding rollout outcomes:

where OO denotes the tool rollout process. We compute a reward ri,tr_{i,t} for each step based on the final answer quality and intermediate tool outputs, and derive group-relative advantages A^i,t\hat{A}_{i,t} within each batch of trajectories. The resulting GRPO objective becomes:

To teach the model in learning when and how to invoke tools, we implement a rule-based accuracy reward to optimize the model. For the ii-th question, we define the terminal reward with ground-truth answer a(i)a^{(i)} and model prediction a^(i)\hat{a}^{(i)}:

Vision Trajectory Construction

With the OpenThinkIMG architecture established, training effective tool-using agents requires high-quality tool-use trajectories. In this section, we propose a novel method to batch-generate trajectory data for solving complex reasoning problems using vision tools. The dataset construction algorithm is presented in Algorithm 1. The process is formally described below in three steps:

For each example (Q(i),I(i))(Q^{(i)},I^{(i)}), we leverage GPT-4o’s few-shot task decomposition capabilities to produce an initial action plan: \rho^{(i)}=\bigl{(}a_{1}^{(i)},a_{2}^{(i)},\dots,a_{n^{(i)}}^{(i)}\bigr{)}, where each at(i)a_{t}^{(i)} is chosen from our predefined vision tools. At this stage, the model performs a symbolic reasoning process to determine the necessary steps without executing any operations. It effectively identifies and schedules the required actions based on its internal understanding of the problem context and the task requirements. To ensure both quality and coherence, we have meticulously designed five demonstration examples to guide the model’s generation process. Moreover, we sample with a moderate temperature (T=0.7T=0.7) to encourage exploration and reject any plans lacking essential steps or containing unsupported actions. The prompt for generating tool-use trajectory is shown in Figure 5.

2 Rationale Parsing and Tool Call Completion

Given the symbolic plan ρ(i)\rho^{(i)}, we batch invoke the corresponding vision tools via our tool server, obtaining rollout outputs: \omega^{(i)}=\bigl{(}o_{1}^{(i)},o_{2}^{(i)},\dots,o_{n^{(i)}}^{(i)}\bigr{)}\;\sim\;O\bigl{(}\cdot\mid\rho^{(i)},Q^{(i)}\bigr{)}. We employ a JSON schema and json.loads to parse each tool’s response, automatically aligning ot(i)o_{t}^{(i)} with at(i)a_{t}^{(i)}. To improve efficiency, outputs are cached and processed in parallel batches of size up to B=128B=128. The final output of this stage is a complete reasoning chain in which each planned action is paired with its corresponding tool result:

It is worth noting that this stage focuses solely on rationale completion, and data filtering is addressed in the next section.

3 Filtering and Rule-Based Validation

To ensure trajectory quality, we apply a multi-stage filtering procedure. First, any τ(i)\tau^{(i)} containing malformed JSON or missing outputs is discarded. Next, we use Qwen2-VL-72B (Wang et al., 2024) alongside rule-based checks (e.g., bounding-box consistency, mask coverage, OCR accuracy) to evaluate both the final answer and intermediate rationale. Next, we apply logical consistency checks and discard any trajectory that does not pass. In addition, human evaluation is incorporated to further confirm the accuracy of the filtered data. By combining automated rule-based filtering with manual verification, our approach ensures that only high-quality reasoning paths are used for training, thereby providing a solid foundation for the Cold-Start and V-ToolRL stages.

Chart Reasoning Experiments

After describing the OpenThinkIMG framework for tool-augmented reasoning and our novel method for constructing vision tool-use trajectories, in this section, we turn to empirical validation on chart reasoning tasks. First, we introduce the data collection process. Next, we describe the vision tools and invocation strategies applied to chart reasoning. We then outline the experimental setup, including training configurations and baseline comparisons. Finally, we present a comprehensive analysis of current performance and outline directions for future work.

For the specific domain of chart reasoning, we strategically selected a subset of the vision tools detailed in Section 2.1. This selection prioritizes capabilities essential for deconstructing graphical data and extracting both quantitative and qualitative insights. Key operations facilitated by these tools include precise spatial localization of data points or chart elements (i.e., using Point), visual annotation to correlate values across axes or highlight thresholds (i.e., via DrawVerticalLineByX and DrawHorizontalLineByY), and focused regional analysis through localized segmentation or magnification (i.e., using SegmentRegionAroundPoint and ZoomInSubfigure). Crucially, robust text extraction (via OCR) is employed to interpret axis labels, legends, titles, and embedded data values. The combined application of these tools, guided by the LVLM’s reasoning, allows for a systematic approach to understanding chart structures and retrieving the visual evidence necessary to answer complex queries.

2 Experimental Setup

We selected the ChartGemma dataset (Masry et al., 2024) because its samples necessitate step-by-step problem-solving, providing an ideal testbed for evaluating our V-ToolRL approach’s ability to learn adaptive tool usage. We partitioned the dataset into a training set of 14,501 samples and a test set of 1,000 samples. To initialize the model’s policy via Cold-Start, we curated a specialized training subset. We generated 1,471 tool-use trajectories using the method from Section 3. To mitigate the risk of the model overfitting to specific tool sequences and to preserve its general reasoning faculties, we augmented this trajectory data with an equivalent volume of text-based CoT reasoning data, also drawn from the training set. This mixed dataset, totaling 2,942 examples, formed the basis for our Cold-Start process. Subsequently, the entire pool of 14,501 training samples was utilized during the V-ToolRL training period, providing the environment for the agent to explore and learn the optimal tool invocation policy.

We conducted model training using configurations with either four or eight NVIDIA Tesla A100 GPUs. To facilitate efficient parallel training, we employed DeepSpeed Zero-Stage 3 (Ren et al., 2021) and FlashAttention-2 (Dao, 2023). The Qwen2-VL-2B-Instruct model served as our main backbone. The training process involved two phases: ❶ Cold-start Period: Models were trained for 2 epochs with a learning rate of 2e-5 and a batch size of 128. We utilized a cosine learning rate scheduler featuring a 3% warm-up period. ❷ V-ToolRL Period: For this phase, models were trained for 500 steps. We employed the AdamW optimizer with an initial learning rate of 1e-6. The maximum sequence length was set to 2048 tokens, the batch size was 144, and the KL divergence coefficient (β\beta) was configured to 0.0.

We compare V-ToolRL against the below models and approaches: ❶ GPT-4.1: OpenAI’s state-of-the-art multimodal model (OpenAI, 2024). Evaluated in a zero-shot setting without external tools as a strong generalist baseline. ❷ Gemini-2.0-flash-exp: Google’s high-capability multimodal model (Gemini Team, 2023). Also evaluated zero-shot without external tools. ❸ Taco: Learns to invoke 15 external tools (e.g., OCR, calculator) by generating Chain-of-Thought-and-Action (CoTA) sequences (Ma et al., 2024b). Taco is trained via supervised learning on synthetic CoTA data and typically executes tools within a single process, contrasting with our RL-based approach and distributed architecture. ❹ CogCom: Employs external visual tools through a Chain of Manipulations (CoM) paradigm for step-by-step reasoning (Wu et al., 2023a). Similar to Taco, it is trained via supervised learning on CoM data and generally integrates tool execution within a unified process, differing from our V-ToolRL method and deployment strategy. Furthermore, to dissect the contributions of our framework’s core components, we evaluate the following internal variations, which also serve as progressive baselines: ❺ Qwen-Base: The foundational Qwen2-VL-2B model without any tool-use fine-tuning or reinforcement learning, representing the starting point. ❻ Qwen-SFT: The Qwen2-VL-2B model after supervised fine-tuning (Cold-Start stage) on our generated tool-use trajectories, quantifying the benefit of initial policy learning. ❼ Text-based RL: An RL agent trained similarly to V-ToolRL but without direct integration of visual tool outputs in its state or reward, isolating the impact of the “visual” component. These internal baselines are crucial for understanding the incremental benefits of each design choice in V-ToolRL.

3 Main Results

We evaluate the performance of our proposed V-ToolRL method against the above baselines. The results, measured in accuracy (%), are presented in Figure 2.

As shown in Figure 2(left), V-ToolRL achieves 59.39% accuracy on the ChartGemma test set. This performance significantly surpasses other open-source tool-augmented frameworks such as Taco-8B (30.50%) and CogCom-13B (15.07%). This advantage is particularly notable given that V-ToolRL utilizes a 2B parameter Qwen2-VL base model, whereas these counterparts employ larger 8B and 13B parameter models. The results strongly suggest that our reinforcement learning paradigm for adaptive tool selection is more effective than supervised methods reliant on predefined CoTA or CoM action sequences. When compared to high-capability closed-source models, V-ToolRL (59.39%) not only demonstrates a marked ability to enhance open-source model performance but also notably outperforms GPT-4.1 (50.71%) and achieves a competitive result when compared to Gemini (68.20%) on these complex chart reasoning tasks requiring structured tool interaction.

Figure 2 (Right) presents an ablation study on the Qwen2-VL-2B backbone. The base model (Qwen-Base) scored 29.56%. Supervised Fine-Tuning with our generated trajectories (Qwen-SFT, representing the Cold-Start stage) improved accuracy to 45.67%, indicating the benefit of initial tool invocation learning. A Text-based RL baseline, using RL without direct visual tool output integration, achieved 51.63%. Our full V-ToolRL framework, which integrates visual feedback from tools into the RL process, attained the highest accuracy at 59.39%. This represents a +29.83 point improvement over the base model and a +13.72 point gain over SFT alone. The +7.76 point advantage of V-ToolRL over Text-based RL specifically highlights the importance of the “V” component. This component, representing the direct integration of visual tool outputs, is crucial for maximizing performance on visually grounded tasks like chart reasoning. These findings confirm the significant contributions of both the Cold-Start initialization and the subsequent vision-integrated V-ToolRL training stages to learning complex, sequential, and adaptive tool invocation strategies.

4 Analysis of Tool Invocation Efficiency

To understand how V-ToolRL influences tool utilization, we tracked the average number of tool calls per sample throughout the training phase, as depicted in Figure 3(a). The plot reveals a distinct learning curve: initially, the average tool usage is relatively high, around 0.63 calls per sample, likely reflecting an early exploratory phase or the initial policy from the Cold-Start stage. However, as training progresses, there is a rapid and substantial decrease in tool invocation (Qu et al., 2025). By approximately 250-300 training steps, the average number of tool calls stabilizes at a remarkably low value, roughly between 0.10 and 0.12 calls per sample. This pronounced downward trend strongly indicates that the reinforcement learning process effectively instills tool-use efficiency. The agent learns to be highly selective, invoking tools primarily when their utility offers a clear path towards maximizing rewards, thereby implicitly penalizing superfluous or redundant tool calls. This outcome demonstrates V-ToolRL’s capability to foster a parsimonious yet adaptive tool invocation strategy, preventing indiscriminate overuse of available tools.

5 Development of Reasoning Complexity

Concurrently with the optimization of tool efficiency, we examined the evolution of the agent’s output complexity by monitoring the average completion length during V-ToolRL training, shown in Figure 3(b). This metric exhibits a clear and consistent upward trajectory. Starting from an initial average length of approximately 66 tokens, the model’s output progressively lengthens, eventually plateauing in the range of 83 to 86 tokens by around 400-450 training steps. This steady increase in completion length suggests that as the agent becomes more proficient in leveraging tools through V-ToolRL, it simultaneously develops the capacity to generate more elaborate and detailed reasoning narratives. These extended completions likely encompass more comprehensive Chain-of-Thought (CoT) steps, explicit justifications for tool usage, and better integration of information derived from tool outputs. Such detailed reasoning is essential for addressing the complexities inherent in tasks like chart analysis, demonstrating that V-ToolRL encourages not just effective tool use but also the generation of thorough and interpretable reasoning paths.

6 Learning Dynamics and Impact of Visual Feedback

Figure 3(c) illustrates the learning dynamics during the V-ToolRL phase by plotting the reward accuracy on the training set against training steps. This subplot directly compares our full V-ToolRL approach (orange curve) with a Text-based RL baseline (blue curve), which lacks direct integration of visual tool outputs. Several key observations emerge: Firstly, V-ToolRL consistently achieves higher reward accuracy throughout the training process, starting from a better initial point and maintaining a significant performance margin over the Text-based RL baseline. Secondly, V-ToolRL exhibits a steeper learning curve, particularly in the initial 100-200 training steps, indicating faster convergence towards effective policies. While both approaches show signs of plateauing towards 500 steps, V-ToolRL stabilizes at a substantially higher accuracy level. This persistent gap underscores the critical contribution of incorporating visual feedback from tool interactions directly into the reinforcement learning loop. The superior performance and learning efficiency of V-ToolRL affirm that enabling the agent to “see” and react to the visual outcomes of its tool use is paramount for mastering complex, visually grounded reasoning tasks.

7 Qualitative Case Studies

Beyond aggregate metrics, Figure 4 illustrates V-ToolRL’s superior reasoning accuracy and interpretability through learned tool invocation, compared to GPT-4.1’s direct visual interpretation. Our framework effectively decomposes complex queries into verifiable, tool-based subtasks. In a pie chart analysis (Top of Figure 4), V-ToolRL uses ZoomInSubfigure and OCR for precise value extraction, correctly calculating a 15.0% difference. GPT-4.1’s direct visual reading, however, misinterprets values, yielding an incorrect 22.0%. This demonstrates the robustness of tool-assisted data extraction for dense charts. Similarly, for a line graph trend analysis (Bottom of Figure 4), our model uses Point and DrawVerticalLineByX to accurately compare intensity changes, correctly identifying a three-way tie. GPT-4.1, lacking these explicit grounding tools, fails to discern this tie. These cases show V-ToolRL’s policy of leveraging tools for targeted information gathering and visual augmentation results in more accurate and transparent reasoning than direct interpretation, especially where precision is critical.

Related Work

LVLMs have rapidly advanced multimodal understanding. Their development began with foundational pre-training on image-caption datasets (Jia et al., 2021; Lin et al., 2014), establishing initial vision-language grounding. Subsequently, sophisticated architectures emerged focusing on effective alignment between visual encoders and powerful LLMs, exemplified by seminal models like Flamingo (Alayrac et al., 2022) and BLIP-2 (Li et al., 2023). A significant leap in capability was achieved through instruction tuning, which enabled models to follow complex visual directives with greater fidelity. This paradigm led to the development of influential model families, notably Qwen-VL-series (Bai et al., 2023; Wang et al., 2024; Bai et al., 2025) and LLaVA-series (Liu et al., 2024b, 2023a). However, current LVLMs often falter on tasks requiring intricate multi-step visual reasoning or precise interaction with visual content (Wu et al., 2023a; Ma et al., 2024b). Cognitively, they tend towards pattern recognition rather than deeper, human-like visual manipulation or scaffolded thought (Zhang and Norman, 1994; Larkin and Simon, 1987). This highlights the necessity of integrating vision tools to enable more granular, verifiable interaction and emulate “thinking with images”. While recent RL-based efforts (e.g., MM-Eureka (Meng et al., 2025), LMM-R1 (Peng et al., 2025)) have focused on enhancing intrinsic reasoning, they typically do not address external tool interaction. To our knowledge, our work is the first to present an end-to-end solution for learning adaptive external vision tool policies in LVLMs.

To address the complexities beyond the reach of standalone LVLMs, augmenting them with external tools is a rapidly growing research area. This allows models to leverage dedicated functions for tasks like OCR, calculation, grounding, or knowledge retrieval. While early methods explored prompting (Wu et al., 2023b), recent focus has shifted to trainable frameworks. Models like LLaVA-plus (Liu et al., 2023b), MLLM-Tool (Wang et al., 2025), TACO (Ma et al., 2024b), and CogCom (Wu et al., 2023a) explicitly train LVLMs for tool interaction, typically via supervised fine-tuning on synthetically generated execution traces (e.g., CoTA or CoM paradigms). Other related works improve specific tool-assisted capabilities like grounding (Liu et al., 2023c) or detailed visual search (Wu and Xie, 2023). Despite these advancements, significant challenges persist. Firstly, the field lacks a unified approach to tool definition and interfacing, with heterogeneous implementations hindering standardization and reproducibility (Ma et al., 2024b). Secondly, reliance on SFT often yields policies with limited adaptability and generalization to novel scenarios (Chu et al., 2025b, a). To address this gap, our work offers distinct solutions. OpenThinkIMG provides a standardized, distributed framework for modular tool deployment, addressing heterogeneity and scalability. Furthermore, our proposed V-ToolRL employs reinforcement learning, enabling agents to learn adaptive tool-use policies that generalize beyond fixed SFT trajectories. To our knowledge, this combination represents a novel end-to-end approach for robust and flexible tool-augmented visual reasoning in LVLMs.

Conclusion

In this work, we addressed the limitations of supervised learning for training LVLMs to dynamically utilize external vision tools. We introduced OpenThinkIMG, a platform designed to standardize tool integration and facilitate the training process, and proposed V-ToolRL, a reinforcement learning framework for learning adaptive tool invocation policies. V-ToolRL enables agents to optimize tool selection and sequencing through direct interaction and reward feedback, moving beyond the constraints of mimicking static trajectories. Our experiments on chart reasoning empirically validated this approach, showing that V-ToolRL significantly improves performance over SFT initialization and outperforms existing supervised tool-learning methods, while fostering efficient tool usage. This demonstrates the efficacy of RL in equipping multimodal agents with robust, interactive reasoning capabilities. We hope that OpenThinkIMG, coupled with the V-ToolRL methodology, will serve as a valuable resource for the community, accelerating research into adaptive multimodal agents capable of sophisticated, interactive visual reasoning.

References

Appendix A Prompts for Synthetic Trajectory Generation

Figure 5 shows the prompt used for generating high-quality tool-use trajectories.