Kwai Keye-VL 1.5 Technical Report

Biao Yang, Bin Wen, Boyang Ding, Changyi Liu, Chenglong Chu, Chengru Song, Chongling Rao, Chuan Yi, Da Li, Dunju Zang, Fan Yang, Guorui Zhou, Guowang Zhang, Han Shen, Hao Peng, Haojie Ding, Hao Wang, Haonan Fan, Hengrui Ju, Jiaming Huang, Jiangxia Cao, Jiankang Chen, Jingyun Hua, Kaibing Chen, Kaiyu Jiang, Kaiyu Tang, Kun Gai, Muhao Wei, Qiang Wang, Ruitao Wang, Sen Na, Shengnan Zhang, Siyang Mao, Sui Huang, Tianke Zhang, Tingting Gao, Wei Chen, Wei Yuan, Xiangyu Wu, Xiao Hu, Xingyu Lu, Yi-Fan Zhang, Yiping Yang, Yulong Chen, Zeyi Lu, Zhenhua Wu, Zhixin Ling, Zhuoran Yang, Ziming Li, Di Xu, Haixuan Gao, Hang Li, Jing Wang, Lejian Ren, Qigen Hu, Qianqian Wang, Shiyao Wang, Xinchen Luo, Yan Li, Yuhang Hu, Zixing Zhang

Introduction

In recent years, Large Language Models (LLMs)(Grattafiori et al. (2024); Abdin et al. (2024); Team (2025a); Wang et al. (2024a)) have experienced rapid development, ushering in a new era of artificial intelligence with their powerful capabilities in understanding (FaceBook (2025); Team (2025b)), generation (Yang et al. (2025); Seed et al. (2025)), and linguistic reasoning (Guo et al. (2025a); Liu et al. (2024a)). This wave has also driven the rapid advancement of Multimodal Large Language Models (MLLMs) OpenAI (2025); Chen et al. (2024a; b); Hurst et al. (2024); Team et al. (2025a); Feng et al. (2024); Fu et al. (2025a); Han et al. (2024); Li et al. (2023); Luo et al. (2023); Guo et al. (2025b); Team et al. (2025b); Zhang et al. (2025a)), which extend powerful language capabilities to the visual domain, enabling the execution of complex tasks such as visual question answering (Li et al. (2024); Chen et al. (2024c)), detailed image description (Luo et al. (2024a); Rang et al. (2025); Li et al. (2025a)), object localization (Bai et al. (2025); Ma et al. (2025)), and visual reasoning (OpenAI (2025); Su et al. (2025); Hu et al. (2025a)).

Despite significant progress in static image understanding, video understanding remains a major challenge. Video content is inherently more dynamic and information-dense than static images, requiring models to process temporal relationships and sequential information while managing the fundamental trade-off between temporal coverage and spatial resolution. Existing approaches typically employ uniform frame sampling under fixed resolution constraints, which leads to suboptimal performance when fine-grained visual details and temporal consistency are required for content understanding (Shen et al. (2025); Lin et al. (2023); Luo et al. (2024b); Team et al. (2025c); Bai et al. (2025)).

To address these limitations, we propose Keye-VL-1.5, an 8-billion parameter multimodal foundation model that achieves state-of-the-art performance in video understanding while maintaining robust capabilities in general vision-language tasks. Our contributions span three key areas: architectural innovations for efficient multimodal processing, progressive pre-training strategies, and comprehensive post-training methodologies.

Architecture and Slow-Fast Video Encoding: We propose a novel Slow-Fast video encoding strategy that dynamically allocates computational resources based on inter-frame similarity. Key frames with significant visual changes are processed through the Slow pathway at higher resolution, while relatively static frames are processed through the Fast pathway at lower resolution but with higher temporal coverage. This adaptive approach, guided by patch-based similarity functions, effectively addresses the trade-off between spatial detail and temporal breadth.

Progressive Pre-training with Long Context Extension: Our pre-training methodology comprises four carefully designed stages that progressively build multimodal capabilities. Beginning with cross-modal alignment and multi-task learning, we systematically extend the model’s context length from 8K to 128K tokens during the annealing phase, enabling it to process longer videos and more complex visual content. This progressive approach ensures stable training while maximizing the utilization of the extended context window to enhance video understanding capabilities. The final model fusion stage combines models trained with different data mixtures to improve robustness and reduce bias.

Post-training for Reasoning and Human Preference Alignment: Our post-training process focuses on two critical aspects: enhancing reasoning capabilities and aligning with human preferences. We develop a comprehensive pipeline with three key components. First, we design a 5-step chain-of-thought reasoning data construction pipeline to generate high-quality cold-start data. Second, we employ the GSPO algorithm for verifiable reward-based reinforcement learning training. This includes progressive prompt sampling to handle difficult samples. Specifically, for samples where the model consistently fails during multiple rollouts, we provide varying levels of hints in the prompt to improve the efficiency of the rollouts. We use the RL model to generate better SFT data, and then perform the next round of RL training based on the SFT model, continuously iterating. Finally, we conduct alignment reinforcement learning training to enhance instruction following, response formatting, and preference alignment. This systematic approach ensures that Keye-VL-1.5 achieves excellent benchmark performance while providing responses that align with human expectations and preferences.

Through evaluation on public benchmarks and rigorous internal human assessment, we validate that Keye-VL-1.5 demonstrates significant improvements compared to existing models, particularly in video understanding tasks. Our work provides practical solutions for building next-generation multimodal models capable of complex video understanding and reasoning.

Model Architecture

Figure 2 gives a high-level overview of our Keye-VL-1.5, which follows a classic MLLM architecture that includes three key components: a Vision Transformer (ViT), a MLP projector, and a language decoder. For ViT component, we apply the open-source SigLIP-400M-384-14 https://huggingface.co/google/siglip-so400m-patch14-384 as our vision encoder to extract vision information. For LLM component, we employ the widely used Qwen3-8B as our language decoder, to provide the universal world semantic knowledge understanding capabilities. For the projector, we randomly initialize its parameters and fully pre-training it at the Stage 1. In the following sections, we provide our key upgrades, data pipeline and training recipes.

In past years, many MLLMs efforts have adopted the well-trained fixed-resolution ViTs as their vision encoders, such as ViT-bigG (Cherti et al. (2023) ), SigLIP-400M (Zhai et al. (2023)) and others. However, unlike pre-trained CLIP-based ViTs (Radford et al. (2021) ) that only handle coarse-grained image-caption matching task during training, MLLMs often tackle various finer-grained generation tasks, existing a large gap between them. Therefore, we anticipate that our ViT will possess the following capabilities: during processing, images and videos maintain their structural integrity and all details are preserved.

To this end, there are some pioneer MLLMs exploring native-resolution ViT in recent years, such as Qwen2.5-VL, Seed-VL-1.5, Kimi-VL, etc. In Keye-VL-1.5, we also implement a native-resolution ViT, to naturally process images at original resolution, avoiding some complex and redundant image splicing/splitting operations (e.g., MiniCPM2 (Yao et al. (2024))). Specifically, our ViT is initialized by the SigLIP-400M-384-14, a fixed-resolution variant with absolute learnable position embeddings to inject the spatial information. According to it, we first employ interpolation techniques to extend fixed-length learnable position embeddings into resolution-adaptive position embeddings, enabling our basic native-resolution modeling while preserving the pretrained workflow. Afterwards, to further enhance extrapolation capabilities for positional encoding along visual dimensions, we introduce 2D Rotary Position Embedding (RoPE) to strengthen the visual information modeling. In our trial experience, we observe that incorporating 2D RoPE significantly improves the model’s performance on high-resolution image. Finally, building upon the two types of position embeddings, we incorporate the NaViT packing with FlashAttention techniques to continue training our ViT across images with varying resolutions.

During the ViT pre-training procedure, we optimize our native-resolution modifications via SigLIP loss function (the text tower is also from SigLIP-400M-384-14). We use the same distribution data as the downstream MLLM for training, including a total of 500B Tokens from open source data DataComp (Gadre et al. (2023)), LAION (Schuhmann et al. (2022)), CC12M (Changpinyo et al. (2021)), PD12M (Meyer et al. (2024)), COCO (Lin et al. (2014)) and other in-house data.

2 Visual Encoding

To guarantee that our language decoder can perceive enough visual signals to understand images and videos in detail, we devise different modeling strategies for them:

Native-Resolution Image Encoding: for images encoding with different resolutions, we set the total number of tokens for each image to 20,480 (at LLM side), which can cover images with more than tens of million pixels and is sufficient to help the model to see the enough details of images.

Slow-Fast Video Encoding: for the video encoding with vary FPS, resolutions and duration, linearly increasing any of these factors would lead to a sharp increase in the token budget on the LLM side, thus making it challenging to strike a balance between performance and cost. To our knowledge, most existing MLLMs typically adopt a fixed number of frames and accordingly reduce the resolution of each frame to meet token budget limitations. Following the paradigm, Qwen-2.5-VL further proposes 2D convolution technique to merge the adjacent frames, aiming to enable the LLM decoder to perceive more video signals within a fixed frame count. Nevertheless, under the uniform frame sampling strategy, although many adjacent frames may be highly similar, there can still be some cases where consecutive frames show significant differences, especially when sampling-interval is larger, a person is moving or viewpoint is shifting. As a result, the rough 2D convolution merging technique maybe unfriendly to effective video understanding, since it relies on overly strong assumptions. Considering the inherent characteristics of video: where adjacent frames are mostly similar yet sometimes significant changes, we propose a SlowFast video encoding strategy:

Slow Pathway: This pathway is designed to capture visual information from rapidly changing frames. It operates at a lower number of frames but with higher resolution.

Fast Pathway: In contrast, the Fast Pathway captures subtle changes visual signal from relatively static frames. It uses a higher number of frames but at a lower resolution.

To identify the slow/fast frames from the video, we first devise a patch-based similarity function to extract them: (1) The first frame is always defined as a slow frame; (2) For each subsequent frame, if its patch similarity with the latest slow frame exceeds 95%, it is marked as a fast frame; otherwise, it is marked as a new slow frame. After obtaining the slow and fast frames, we set the fast frame’s token budget to 30% of a slow frame’s budget to balance the trade-off between frame numbers and the total token budget. Then, we utilize a binary search technique to precisely calculate the number of tokens per slow frame under the total token budget limitation (e.g., 75,000 tokens in Keye-VL-1.5). Meanwhile, to more clearly identify the boundaries and timestamp information between Slow and Fast frames, we introduce additional special tokens along with absolute timestamps to guide the model during learning, as shown in Figure 3.

Pre-Training

In this section, we first describe the construction of the pre-training dataset, followed by an overview of the overall training pipeline and configuration.

In our data construction pipeline, we have assembled a diverse, high-quality corpus with exceeding 1 trillion tokens to support our models training, sourced from both public datasets and proprietary in-house data. Generally, our training data encompasses six primary categories: Image Caption, OCR & VQA, Grounding & Counting, Interleaved, Video Understanding and Pure Text data. To ensure these overall data quality, we have designed customized filtering mechanisms tailored to the characteristics of each data category. For large volumes of medium-quality data, we employ CLIP (Radford et al. (2021)) scores for preliminary filtering. For smaller amounts of high-quality data, we utilize open-source MLLMs as discriminators for data selection. Additionally, we also conduct rigorous image-based deduplication operation, to avoid the potential data leakage between our training corpus and evaluation benchmarks (Dixit et al. (2021)). Specifically, we identify highly similar images, then remove these near-duplicates from the dataset. In the following sections, we provide detailed descriptions of each category of our data.

Image caption task provides the fundamental world knowledge to establish a mapping relationship between visual features and linguistic concepts by pairing image with textual descriptions. Based on large-scale caption data, our model gains the ability to perceive and comprehend a broad, rich spectrum of world knowledge, such as real-world physical principles and cultural conventions. Although we can public access many diverse Chinese and English open-source caption data source, such as LAION (Schuhmann et al. (2022)), DataComp (Gadre et al. (2023)) and Coyo (Byeon et al. (2022)), the quality of such data is often unreliable, as it typically only undergoes simple crawler-based matching.

To alleviate such data noise, we conduct strict similarity-based filtering pipeline to control the data quality, e.g., scoring the raw rigorous image-caption pair by a CLIP model. In practice, to ensure data quality, we retain high-similarity image-caption pairs (e.g., CLIP score ¿ 0.9) while leveraging filtered low-quality open-source image data and our in-house image data through a re-captioning pipeline. During the re-caption, we utilize several MLLMs (Qwen2.5-VL 72B (Bai et al. (2025)), Tarsier2 (Yuan et al. (2025)), GPT-4o (Hurst et al. (2024)), Gemini1.5-pro (Team et al. (2023)) and others) to generate the synthesis caption for vary resolution images and image category information. In our experience, we find that recaption data generated by different MLLMs can be very helpful for fine-grained image understanding.

Further, to avoid our model degenerate into a caption generators and hurt its instruction-following and complex reasoning abilities. We implemented a data augmentation strategy with multiple-caption/question-answering pair to maintain our model’s general conversation and instruction capabilities:

¡image, caption, [eos], question, answer¿ format data: training our model to seamlessly transition from generating captions to accurately answering follow-up questions, thereby strengthening contextual understanding and task continuity.

¡image, question, answer, [eos], caption¿ format data: Reverses the task order, requiring the model to answer before describing, which helps break the tendency to default to caption generation and improves task-switching flexibility and instruction sensitivity.

Instruction-following image captioning/QA: We first provide dozens of images as input, then randomly ask questions or generate captions corresponding to specific images.

Besides, to improve our model robustness and faithfulness, we proactively inject some ‘trap questions’ that refer to non-existent or contradictory questions. These counterfactual data would encourage the model to ground its responses more accurately in visual content rather than textual priors.

1.2 OCR & VQA Data

Optical Character Recognition (OCR) and Visual Question Answering (VQA) are vital tasks to encourage our model to distinguish the details of images. By integrating OCR capabilities, the model can accurately extract and interpret textual information within images, while VQA task enables our model to comprehend and reason about visual content in a context-aware manner. In order to build our capabilities in OCR and VQA, we have collected a large number of open-source data, such as Latex-Formula, hand-write text, real-world street views, charts, rich-text documents, multi-image OCR and so on. Since most of the open-source datasets are in English, to further enhance the model’s capability in Chinese OCR & VQA tasks, we introduce multiple techniques for synthesizing in-house Chinese data:

Synthesis with SOTA MLLMs: To enhance the model’s OCR capabilities, we extract images from both open-source and in-house image-text datasets to build our image repository, utilizing the text-dense images from it to synthesize comprehensive OCR dataset which covering diverse scenario. For VQA task, we first design a set of seed-questions and expand the initial question pool through self-evolution methods. Next, both images and their corresponding captions are fed into SOTA MLLMs to generate high-quality and diverse VQA data, such as utilizing Qwen2.5-VL-72B to generate multi-turn challenging question-answering pairs.

Rendering with Font Tools: Considering the scarcity of high-quality open-source Chinese OCR data, we further leverage font rendering tools to synthesize high-quality OCR samples (includes (1) diverse image backgrounds/layout, (2) semantic/non-semantic text, (3) multiple fonts styles/sizes and (4) vary image resolutions), which significantly enhances the model’s robustness for Chinese OCR recognition.

Structured Document and Code Understanding: Further, we also perform complex text recognition tasks by using a vast codebase (e.g., Markdown, HTML, and other programming languages). By rendering codes/documents that preserve their original layout, we could create elaborate OCR tasks, such as reconstructing source code from an image or completing missing code at specific locations, thereby training the model to understand textual hierarchy and structure.

Instruction Following OCR: Moreover, to enhance our model’s capability to follow specific OCR instructions (e.g., “extract only the text from the third column”), we built an instruction-following OCR dataset. Each sample consists of a character-matrix image paired with a Chinese instruction, covering thirteen classes of reading and locating templates (e.g., row/column extraction in four directions and their combinations). This dataset is enriched with diverse text sources, noise injection (English, Japanese, Korean, symbols, numbers, uncommon characters, and emojis).

1.3 Grounding & Counting Data

Object grounding is one of the fundamental abilities of MLLMs( Bai et al. (2025); Seed et al. (2025)), which enables our model to establish a direct connection between temporal/visual information and text semantics, as shown in the Table 1. In Keye-VL-1.5 objective grounding, we primarily utilize three object localization forms: center points, bounding boxes, and polygons. Their coordinates are strictly typed as integers and normalized to the range [0, 1000) for different resolution images, . In general, we mainly employ the RefCoCo (Kazemzadeh et al. (2014)), VisualGenome (Krishna et al. (2017)), TolokaVQA (Ustalov et al. (2023)) as our grounding data source, and the PixMo (Deitke et al. (2024)) as our counting data source. For the in-house grounding data generation, we use other MLLMs (e.g., Gemini 2.5 Pro, Qwen-2.5-72B) to extract the answer area bounding boxes of corresponding document questions. To filter the incorrect, missing, or ambiguous annotation grounding data, we utilize the CLIP and Qwen-2.5-7B to select the higher-score points/boxes/polygons as our training data, i.e., extracting the corresponding grounding area from the image to compute its similarity with the target objective text.

For temporal grounding data, we construct a three-step coarse-to-fine-grained data synthesis pipeline based on our massive short-videos base. In the first step, we employ the TEMPURA Cheng et al. (2025a) to process a given short-video as several event video clips with their temporal captions. Next, to alleviate the “repetitive collapse” in redundant or meaningless descriptions issue of raw TEMPURA outputs, we apply the SOTA MLLMs as a filter to identify and remove such low-quality, repetitive event video clips to obtain reliable temporal grounding captions. At last, according to those captions, we further utilize the Gemini 2.5 Pro to enrich our database to generate a series of logical question-answering pairs about timestamps, which could empower our model’s understanding of temporal causality relationships. In this way, our pipeline ensures our model not only describes what happens in a video, but also understands and reasons about when and why.

1.4 Interleaved Text-Image Data

Instead of the learning task surrounding the single images, we also introduce a large amount of interleaved data to enhance our language decoder’s longer multi-modal context modeling ability and longer sequence adaptation, e.g., 128K context modeling. Actually, beyond modeling multi-image correlations, the interleaved data could contribute several critical advantages in pre-training: (1) Preservation of General Knowledge: It contains a wealth of universal knowledge, ensuring that the LLM module’s core capabilities are not degraded during training, (2) Enhanced Vision-Language Alignment: By leveraging in-context learning, it helps the model better align visual and semantic signals in language model side, (3) Improved Generalization: The diverse and interleaved nature of the data strengthens the model’s ability to reason across modalities and generalize to unseen tasks. Besides the open-source interleaved data, we also build a large-scale in-house interleaved data generation pipeline. Specifically, we focus on the two type of raw rich-text documents processing, the academic PDF data and structured knowledge data, especially the Science, Technology, Engineering, and Mathematics (STEM) data. We collect a substantial amount of academic and knowledge-based PDF/structured data to render the text content into plain text format and insert the corresponding images at their original positions within the text. In such a process, we conduct rigorous data protection strategies to ensure high-quality outputs. Our pipeline includes: (1) Garbled character recognition: identifying and removing garbled characters, (2) Low-resolution/broken image filtering: ensuring image quality, (3) Text-image similarity validation: ensuring semantic alignment between interleaved image-text.

1.5 Video Data

As a short-video and live-streaming service provider, the video understanding ability is the most important point of Kwai, such as understanding the video details, generating summaries, and expressing interesting implications. To reach the goal, our video data are collected from multiple sources, including diverse open-source datasets (ShareGPT4V, Pandas and others) and a large-scale high-quality in-house video data. Based on these videos, we conduct the following key pipelines to guarantee our data quality:

Interleaved video-ASR: For audio signals, we currently use speech-to-text tools (e.g., Qwen2.5-Omni (Xu et al. (2025a))) to recognize them, and then form a interleaved style to connect images and audio to our model.

Video recaption: With (optional) ASR results, we next utilize diverse public MLLMs to generate its caption under different FPS setting, such as 0.5/1/2.

Frame-level OCR annotation: In order to ensure that our model does not miss any details in each frame, we further added a frame-level OCR task.

In addition to OCR and video captioning/QA tasks, we have designed a series of reasoning-enhanced tasks to help the model better understand contextual relationships in short videos. These include:

Frame-level re-ordering: Given a set of shuffled video frames, our model is required to predict their original chronological order, which enhances its ability to grasp temporal progression and logical flow.

Multiple video matching: Provided with a group of related videos and a set of candidate videos, our model is required to identify the most contextually relevant candidate, which refines its understanding of semantic connections across different videos.

2 Training Recipe

We employ a four-stage progressive training strategy to build a powerful multi-modal foundation model with strong vision-language alignment capabilities. The training pipeline, illustrated in Figure 4, is meticulously designed to ensure that each stage has a clear and interconnected objective.

The Vision Transformer (Dosovitskiy et al. (2020)) (ViT) is initialized with weights from the siglip-so400m-patch14-384 model and undergoes continuous pre-training using the SigLIP (Zhai et al. (2023)) contrastive loss function. This stage focuses on adapting the vision encoder to our internal data distribution. We incorporate native dynamic resolution processing (akin to NaViT (Dehghani et al. (2023))), which preserves the original aspect ratio of images to the greatest extent possible. Additionally, 2D Rotary Position Embeddings (Su et al. (2024)) (RoPE) are integrated to enhance the model’s extrapolation capabilities when processing images of varying resolutions.

The language model is initialized from Qwen3-8B (Yang et al. (2025)). During this stage, the parameters of both the vision and language models are frozen. Training is focused on optimizing the projection MLP layer. With large-scale datasets, we establish a robust alignment between cross-modal features, laying the groundwork for the subsequent learning phase.

All model parameters are unfrozen for end-to-end optimization using a diverse set of multi-task training data. The data in this stage encompasses a wide range of common vision-language tasks, including Image Captioning, Optical Character Recognition (OCR), Grounding, Visual Question Answering (VQA), and interleaved image-text data. This process significantly enhances the model’s fundamental visual understanding capabilities.

This stage involves an annealing phase where the model is fine-tuned on a curated set of high-quality data. The primary goal is to address the issue of insufficient exposure to high-quality samples during the large-scale, broader training of Stage 2. Through optimized learning strategies and data mixtures, we further refine the model’s nuanced understanding and capabilities.

In Stage 1 and Stage 2, we limit the sequence length of each sample to 8,192 (8K), where Data Parallelism is adopted to effectively create large batch sizes. Zero-2 optimization strategy is applied to reduce memory overhead. In the final annealing stage, we extend the context length of the model from 8,192 (8K) to 131,072 (128K). The RoPE inverse frequency of LLM side is reset from 1,000,000 to 8,000,000. The training data is concurrently enriched with high-quality long-context modalities, including long videos, long texts, and large-scale images. Additionally, we switch optimization strategy to Zero-1 and adopt Context Parallelism and Pipeline Parallelism to support long-context training. Under the 128K context length, our controlled experiments show that allocating 24% of tokens to videos, 50% to images, and the remaining 26% to text strikes a good balance between visual capabilities (image and video understanding) and text capabilities.

Post-Training

The SFT data candidate pool contains over 7.5 million multimodal QA samples. We employ the following construction methods to balance comprehensiveness and data quality.

To ensure task diversity, we utilize the proprietary TaskGalaxy (Chen et al. (2025)) framework, which categorizes data across a comprehensive system of 70,000 distinct multimodal task types. We further construct a large amount of data for image/video grounding, counting, GUI, and multi-turn dialogue.

To ensure the data’s challenge, MLLMs are employed to generate multiple reasoning paths for each data point. The complexity of each sample is then measured based on the correctness and length of these responses, allowing for the filtration of overly simple data. We increase the proportion of mathematical, logical reasoning, complex tasks, and long-context data.

To ensure data reliability, human annotators have meticulously crafted captions for the images and videos within the training set.

The training strategy involves a dynamic learning rate. In the later phases of training, the model undergoes an annealing process at a lower learning rate. Evaluations show this annealing step contributes approximately a 1% performance improvement across both open-source and internal benchmarks.

Following SFT, the model undergoes MPO to continuously refine its performance. The MPO dataset includes 250k open-source samples Wang et al. (2024b), 150k text-only samples, and 26k human-annotated samples Zhang et al. (2025b). We perform multiple samplings using Keye-VL-1.5 on the above dataset, and construct multiple pairs of high-quality and low-quality samples using the reward model scores and human annotations. The training strategy for this stage applies the MPO algorithm, utilizing the constructed paired preference data to optimize Keye-VL-1.5’s overall performance.

2 Keye-Reward Model

Recognizing the importance of reward modeling for data quality evaluation and model training, we train our reward model based on the Keye-VL-preview for data filtering and reinforcement learning training. We adapt the Keye-VL-preview model to the reward modeling task with the SFT+RL training process.

Data format: The model input consists of the query, response A, and response B, along with the task definition guiding the model in evaluating the quality of response A and response B. Similar to Keye-VL-1.5’s mix reasoning mode, our training data is composed of two formats: think and no_think. For the no_think mode, the model directly outputs the final judgment based on the input information. In the think mode, the model needs to evaluate the quality of response A and response B separately according to the predefined nine dimensions (such as Credibility, Correctness, Redundancy, Relevance, etc.), then generate a comprehensive evaluation. The mix reasoning mode enables our reward model to reason in terms of efficiency, accuracy, and interpretability.

SFT recipe: The SFT data includes open sourced preference datasets R1-Reward (Zhang et al. (2025c)), MMPR (Wang et al. (2024b)), and manually labeled Keye-VL-preview sampling results. After SFT, we apply data where good responses are shorter than bad responses for annealing to overcome the reward model’s preference for longer responses.

RL recipe: The RL data includes preference data consisting of wrong cases from Keye-VL-preview in the SFT dataset and right cases generated by larger MLLMs, as well as data from MMPR. In this stage, we carefully filter out data with excessively large length differences between positive and negative samples, using format reward and outcome reward as training signals.

We take our reward model to evaluate the quality of Keye-VL’s sampling results, which are applied to update the training data and provide reward signals.

3 LongCoT Cold-Start

After large-scale SFT and MPO, we construct high-quality Long Chain-of-Thought (LongCoT) data for cold-start reasoning training, aiming to enhance Keye-VL’s long CoT reasoning ability, serving as the starting point for subsequent reinforcement learning.

To address the challenge of acquiring high-quality training data for cold-start, we propose a comprehensive five-step automated pipeline for generating LongCoT data, as illustrated in Figure 6. Our approach strategically leverages existing MLLMs to create diverse, high-quality reasoning chains while maintaining both scalability and cost-effectiveness. The pipeline systematically integrates automated generation, rigorous quality assessment, targeted human enhancement, and adaptive data utilization to ensure optimal training data quality across diverse domains and reasoning complexity levels.

Multi-Source Data Collection and Enhancement: Our data generation process begins with the systematic collection of multimodal QA data spanning multiple challenging domains. These domains include mathematical reasoning problems, STEM, OCR and document understanding tasks, visual grounding and object localization, counting, GUI scenarios, and domain-specific business applications. This comprehensive coverage ensures that our generated dataset captures the full spectrum of multimodal reasoning capabilities required for practical applications.

To enhance the complexity and diversity of the collected data, we employ proprietary MLLMs to perform sophisticated question rewriting and task merging operations. The rewriting process transforms simple, straightforward questions into more challenging variants that require deeper reasoning and multi-step problem solving. Additionally, we systematically combine related sub-tasks into comprehensive multi-task instructions, creating scenarios where models must demonstrate proficiency across multiple capabilities simultaneously. This enhancement strategy significantly increases the pedagogical value of each training sample while maintaining natural question flow and coherence.

Multi-Path Reasoning Generation with Confidence Quantification: For each enhanced QA pair, we generate multiple reasoning trajectories leveraging existing MLLMs. A pivotal component of our generation pipeline is the systematic extraction and quantification of model confidence at both the step-wise and holistic response levels. We compute granular confidence scores that capture the model’s certainty in individual reasoning steps as well as the final answer. This confidence metadata serves as a crucial signal for downstream quality assessment and sample prioritization workflows, enabling us to systematically identify the most reliable and coherent reasoning chains from the generated candidate pool. Throughout the multi-round sampling process, we strategically select samples that exhibit diverse logical pathways while maintaining correctness, thereby enriching the diversity of reasoning patterns. Simultaneously, we implement a confidence-prioritized selection strategy, systematically favoring reasoning chains with higher logit-based confidence scores to optimize training sample quality.

Comprehensive Two-Level Quality Assessment: We implement a rigorous two-level quality assessment framework using proprietary MLLMs. This dual assessment strategy operates simultaneously on both answer correctness and reasoning process validity. At the answer level, our assessment framework incorporates flexible matching patterns specifically tailored to different task types and domains. The system supports sophisticated fuzzy matching capabilities and equivalent expression recognition, accommodating variations in phrasing, mathematical notation, and unit representations. For instance, mathematical answers are evaluated considering formula equivalence and unit conversion, while text-based responses account for semantic similarity and paraphrasing.

At the reasoning level, we conduct a granular step-by-step evaluation for each reasoning chain. Every individual reasoning step undergoes scrutiny for logical consistency with preceding steps, factual accuracy against established knowledge, and relevance to the original question. This meticulous evaluation process identifies not only outright errors but also subtle issues such as logical gaps, unsupported assumptions, and irrelevant tangential reasoning. Based on the comprehensive dual assessment results, we categorize all generated samples into three distinct quality tiers.

Category A (High Quality): Both answer and reasoning process are correct.

Category B (Moderate Quality): Correct final answers with reasoning process issues.

Category C (Low Quality): Incorrect answers or severely flawed reasoning, automatically discarded.

Human-in-the-Loop Quality Enhancement: For Category B samples and potentially redundant Category A samples, we implement a systematic human-guided refinement process designed to enhance reasoning quality while preserving valuable training data. Our comprehensive human review protocol encompasses several critical enhancement dimensions:

Category B Sample Refinement: We focus on correcting and streamlining verbose or redundant reasoning steps to improve logical coherence and conciseness. This involves identifying extraneous reasoning chains, consolidating repetitive logical steps, and enhancing the overall flow of argumentation.

Borderline Category A Sample Enhancement: We systematically address samples identified with intermediate redundancy scores during the automated assessment pipeline—those falling below the automatic removal threshold yet still exhibiting suboptimal reasoning patterns.

This human-in-the-loop approach ensures that samples falling into intermediate quality categories undergo systematic improvement rather than wholesale discarding. This methodology strikes an optimal balance between data preservation and quality assurance, thereby enhancing the overall effectiveness of our reasoning dataset for downstream model training.

Dynamic Quality Scoring and Data Utilization Strategy: To optimize data utilization, we implement a comprehensive five-point quality scoring system that evaluates samples across multiple dimensions:

Score 1 (Poor): Simple or ambiguous questions answerable without visual input.

Score 2 (Below Average): Questions with obvious answers or excessive reliance on common sense.

Score 3 (Average): Clear questions requiring basic image understanding but minimal reasoning

Score 4 (Good): Questions demanding reasoning about spatial relationships, etc.

Score 5 (Excellent): Highly multimodal-dependent questions requiring advanced reasoning such as causal inference, occlusion reasoning, or detailed attribute analysis

Based on these quality scores, we implement an adaptive data utilization strategy where higher-quality samples are used more frequently during training. Specifically, samples scoring four or five points are repeated multiple times in the training dataset to reinforce high-quality reasoning patterns, while lower-scoring samples are used sparingly to avoid reinforcing suboptimal behaviors. This strategic approach ensures that the model’s learning process is dominated by the most valuable and challenging examples while maintaining overall dataset diversity.

The entire automated pipeline demonstrates remarkable efficiency and consistency, processing large volumes of input data while maintaining stringent quality standards across diverse domains and task types. The systematic integration of automated generation, rigorous quality assessment, targeted human enhancement, and adaptive utilization creates a comprehensive framework for producing high-quality training data suitable for effective multimodal model cold-start scenarios.

3.2 Model Merging with Domain Specific Experts

We conduct a comprehensive analysis of the LongCoT cold start model’s performance across various benchmarks using the aforementioned training data, with the objective of identifying and addressing model deficiencies prior to the RL phase. Our analysis reveals concentrated weaknesses in three primary domains: pure text processing, mathematical reasoning, and OCR. To address these limitations, we develop a systematic approach involving specialized data collection and expert model training, followed by model merging to enhance Keye-VL-1.5’s foundational capabilities.

OCR Capability Enhancement: Beyond standard OCR datasets, we address specific weaknesses in specialized recognition tasks including license plates, street signage, and official seals. Our enhancement strategy involves three key components: First, we systematically gather OCR datasets targeting identified weak areas, ensuring annotation accuracy through rigorous quality control processes. Second, we develop an automated data pipeline that utilizes images paired with verified OCR annotations to generate relevant OCR questions through other MLLMs, with original annotations serving as ground truth answers to guarantee correctness. Finally, we conduct SFT on the cold-start model using both general-purpose OCR data and our specialized weak-area datasets to create an OCR expert model.

Model Merging: We employ model merging (Li et al., 2025b; Wei et al., 2025) to integrate domain-specific expert models and the LongCoT cold start model into a general model for enhanced performance.

4 Iterative General RL

Based on the cold-start model, we design our General RL process to further enhance Keye-VL-1.5’s reasoning ability, which applies the GSPO (Zheng et al., 2025) (Group Sequence Policy Optimization) algorithm for RLVR (Reinforcement Learning with Verifiable Rewards) training, and employs a cyclical iterative approach to collaboratively enhance both the RL model and the cold-start model.

Training data: We select data from domains including mathematics, science & technology problem, logical reasoning & puzzle problems, code, chart question answering, visual grounding, spatial relationships, and counting to construct the RLVR training set. Each data point contains a verifiable answer used for rule-based reward calculation. We sample data from different domains according to ablation experiments, analyzing the impact of domain-specific data on model metrics. We then increase the proportion of data from domains that contribute positively to performance improvements.

Training Algorithm: Based on sequence-level importance weight, GSPO employs the following sequence-level optimization objective:

where the group-based advantage estimation is defined as:

and the importance ratio based on sequence likelihood si(θ)s_{i}(\theta) is defined as:

4.2 Progressive Hint Sampling

During the training process, we find that the model struggles to generate correct responses for some difficult samples, reflecting a deficiency in the model’s capabilities. To make full use of these challenging samples and enhance Keye-VL-1.5’s reasoning ability, we apply the progressive hint sampling method to improve the success rate of sampling difficult samples.

We first identify the hard cases in the RLVR dataset where Keye-VL-1.5 consistently fails across multiple attempts, then select data with reliable reference answers, sufficient difficulty, and appropriate challenge level as samples for progressive hint sampling.

Unlike the approach of partitioning hints by step, we follow the Minimal Intervention principle to design a hierarchical hint system, aiming to provide the model with the minimal information necessary to solve the problem. We divide the hints into five levels, from abstract concepts to specific reasoning steps:

Level 1 (Concept / Observation): Guide the model to focus on the core concept of the problem or the key features of the image. This level should not contain any problem-solving methods or formulas.

Level 2 (Strategy / Method): Suggest one or more possible problem-solving strategies or approaches. For example, ”Think holistically”, ”Try discussing by cases”, or ”Establish a coordinate system”. This level should not mention specific formulas or calculation steps.

Level 3 (Tools / Formula): Provide hints for specific mathematical theorems, formulas, or tools needed to solve the problem. For example, ”You may need to use the Pythagorean theorem” or ”Consider using integration”. This level should not provide specific calculation steps.

Level 4 (Steps / Calculation): Provide the first concrete operational step in the problem-solving process.

Level 5 (Complete Solution): Provide a complete and clear final solution, a perfect solution that can be used as a standard answer.

For each hard case, we place the hint information after the query and progressively provide hints from low level to high level. When Keye-VL-1.5 can generate correct response based on a particular level of hint, we consider the hint at that level as the minimal information required to help Keye-VL-1.5 solve the hard case. The responses generated based on this minimal information is then applied to update the policy. In Table 9, we report the impact of different levels of hints on the sampling success rate of Keye-VL-1.5 in hard cases, aiming to demonstrate the rationality of our hierarchical hint system and the effectiveness of hints in improving the utilization efficiency of hard cases.

4.3 Iterative General RL & Cold-Start Enhancement

To improve the learning efficiency on reasoning data and break through the performance bottleneck of the SFT model, we design a multi-round iterative paradigm that collaboratively enhances both the cold-start model, which serves as the starting point for General RL, and the model after General RL.

Apply the cold-start model as the initial model and perform General RL training.

Apply the model after General RL for rejection sampling on the Cold Start dataset, score the samples with our reward model. If the sampled results are better than the ground truth, update that data point by replacing the ground truth with the sampled results.

Take the updated cold-start data to train a new cold-start model, which serves as the initial model for the next round of General RL.

Take the updated cold-start model to filter the General RL dataset, selecting data with sampling accuracy between 0 and 1 for next round General RL training.

5 Alignment RL

After General RL, we perform Alignment RL to comprehensively improve the Keye-VL-1.5’s performance in real-world application scenarios. We have developed a diversified task system and reward modeling framework to enhance the model’s capabilities in the following dimensions:

Instruction Following: Improve the model’s ability to generate responses that meet user requirements in terms of content, format, length, and structured output.

Format Adherence: Ensure that the model’s responses conform to predefined formats, such as think-answer, agentic think, auto-think, and no-think.

Preference Alignment: For open-ended questions, enhance the reliability, interactivity, and style of the model’s responses to improve user experience.

The reward system we employ is composed of three main categories:

Rule-Based Reward: Rule-Based reward checks whether the model response adheres to predefined structural and formatting rules, including logical reasoning format (such as think/no_think/auto_think formats), as well as structure-specific guidelines such as json, markdown, and code formatting.

Generative Reward: For data with ground truth that can not be easily evaluated by rules, we design instructions to prompt MLLMs to access model’s response based on how well it aligns with the reference, its reasoning consistency, and the relatedness to key attributes. Additionally, for “security and ethics” tasks, instructions are designed to evaluate whether the responses contain politically errors, misinformation, or offensive content.

Model-Based Reward: For tasks without ground truth, the model’s responses are scored based on our reward model. Our reward model evaluates whether the responses align with human preferences, promoting responses that adhere to ethical standards.

This reward system helps guide the model towards producing accurate, ethical, and contextually appropriate outputs across various tasks.

5.2 Data Construction

For instruction-following task, we design 25 types of hard constraints, including “keywords inclusion,” “punctuation,” “pronunciation,” “output format,” etc., as well as 20 types of soft constraints, such as text style and semantics. We construct a query set consisting of 17k multimodal data and 23k pure text data, with each query assigned 2 to 6 types of constraints as inputs. Hard and soft constraints are rewarded through rule-based rewards and generative rewards, respectively.

For reasoning task, we construct 12k mathematical and logical reasoning queries, with 3 to 5 problem-solving steps designed for each query. The model is required to solve the problem following the prescribed steps. We use rule-based rewards to calculate the correctness of the outcome, and generative rewards to assess whether the reasoning process follows the predefined steps.

For RAG task, we collect a series of instances based on the latest news that require internet searches to obtain answers. We encourage the model to use search and summary behaviors during the think process, ultimately generating the correct answer. We take generative rewards to evaluate the effectiveness of the search behavior in resolving the query, the correctness of the summary behavior, and the consistency of the final answer. We still take GSPO algorithm to optimize our model during Alignment RL.

Training Infrastructure

To efficiently train MLLMs, we make in-depth infrastructure optimization to address three major challenges: architectural heterogeneity, load imbalance, and I/O bottlenecks.

Heterogeneous Hybrid Parallel Strategy: The training bottleneck of MLLMs stems from computational imbalance caused by architectural heterogeneity. The computational characteristics and resource demands of ViT and LLM are vastly different, and unified parallel strategy leads to significant resource wastage. To address this, we design a heterogeneous hybrid parallel strategy: for the relatively fixed computational pattern of the ViT component, we only use data parallelism (DP) to maximize throughput; whereas for the highly parameter- and memory-intensive LLM, we adopt a hybrid parallelism strategy that combines pipeline (PP), tensor (TP), and data parallelism (DP). This refined strategy is a decisive technical prerequisite for achieving 128K ultra-long sequence training of Keye-VL-1.5.

Dynamic Load Balancing Mechanism: Multimodal data inherently leads to load imbalance, primarily due to the correlation between computational load in the visual encoding phase and the input samples. For instance, processing a high-resolution video incurs significantly more computational cost than a static image. In data parallel training, this leads to GPUs processing complex visual input consumes a longer time while other GPUs finish earlier and waits. To address this, we pre-estimate the time complexity of each sample and then use a greedy algorithm to allocate the samples across different GPUs, thereby balancing the total step duration across all GPUs and improving overall hardware utilization.

Flexible and Scalable Dataloader: To fundamentally resolve I/O bottlenecks, we design a flexible and scalable dataloader that deeply senses the topology of parallel training. In terms of data parallelism (DP), each process only loads a shard of the global dataset; in terms of pipeline parallelism (PP), only the first stage (PP0) is responsible for data acquisition and preprocessing; and in tensor parallelism (TP/CP), the data is first fetched by a single process within the group and efficiently broad-casted across processes. Furthermore, we implement an I/O server architecture to offload CPU-intensive tasks such as video decoding from the training nodes, effectively resolving CPU bottlenecks caused by complex media processing. Finally, we implement a instance-level perfect resume mechanism, ensuring that tasks can seamlessly resume from the last successfully processed sample after an interruption, significantly improving the stability and efficiency of large-scale training.

Evaluation

To validate that our continue trained native-resolution ViT is able to capture promising visual representations, we conduct a wide-used zero-shot image classification benchmark analysis. In our evaluation, we perform a comparative analysis between the base SigLIP model and its two native-resolution position embedding variants, leveraging the CLIP Benchmarkhttps://github.com/LAION-AI/CLIP_benchmark framework with text prompt templatehttps://colab.research.google.com/github/openai/clip/blob/master/notebooks/Prompt_Engineering_for_ImageNet.ipynb#scrollTo=sRqDoz1Gbsii.

The evaluation covers six benchmark datasets: ImageNet-1K, ImageNet-V2, ImageNet-A, ImageNet-R, ImageNet-S and ObjectNet, and its results are shown in Table 2. From it, we have the following observations: (1) Compared with base SigLIP model, our 1D interpolation position embedding native-resolution model variant has slightly performance degeneration, the reason might be the interpolated 1D position encoding cannot uniquely identify the underlying 2D patch arrangement. For instance, a sequence of 196 patches may correspond to multiple distinct spatial configurations (e.g., 14×14, 7×28, or 28×7), leading to ambiguous spatial localization during feature projection. (2) With 2D RoPE modification, our ViT could clearly perceive the shape of the image, and showing competitive results with Base SigLIP performance (the best and runner-up results). We think the reason maybe our continued pretraining corpus sharing the same distribution with our MLLMs, rather than the Image-Text matching task.

2 SlowFast Video Encoding Strategy Discussion

In this section, to verify that our SlowFast strategy can capture fine-grained video information, we conduct a comparative analysis between Keye-VL-1.5-Base and Qwen-2.5-VL. Keye-VL-1.5-Base is a pre-trained model equipped with our SlowFast technique, while Qwen-2.5-VL employs a 2D convolution merging technique for video compression.

For a fair comparison, we evaluate both models on the VideoMME benchmark under different settings. Specifically, we test with fixed frame numbers ranging from 32, 64, 128, up to 768, and FPS values from 1 to 4 in increments of 1. Meanwhile, different with linear token budget increasing of 2D convolution along with frame amount, our slowFast strategy has highly adaptive token budget for different videos with different information density. Combining the two factors, we show the prediction performances and the LLM-side visual token budgets across different video category (i.e., short/medium/long and overall) at the Figure 7. According to it, we have the following observations:

In terms of the overall score of VideoMME, our Keye-VL-1.5-Base shows the similar performance trend as Qwen-2.5-VL. Specifically, both of them show an increase then decline performance trend, while our Keye-VL-1.5-Base achieves its best performance at 384 frames, Qwen reaches its peak at 128 frames, indicating that our SlowFast is also a reliable video encoding strategy.

For the sub-category performance of VideoMME, the Qwen-2.5-VL show the inflection point at 128/384/128 and our Keye-VL-1.5-Base shows the inflection point at 192/512/384 for the three categories of short, medium, and long videos. Compared with Qwen-2.5-VL, The fact that Keye’s performance begins to decline at a later point demonstrates that our SlowFast strategy enables the LLM to integrate multi-frame information more effectively.

In terms of video token usage on the LLM side, Qwen-2.5-VL demonstrates a nearly linear relationship with the number of frames and token budget. In contrast, our Keye-VL-1.5-Base generates more visual tokens when the frame number is low, but fewer visual tokens when the frame number is high compared to Qwen-2.5-VL. This phenomenon demonstrates that our SlowFast strategy is more flexible and makes more efficient use of computational resources.

For the different FPS setting, we could observe that our Keye-VL-1.5-Base is more stable in evaluation results than Qwen-2.5-VL. Additionally, our SlowFast video encoding approach achieves a flexible token budget on par with the 2D convolution technique.

3 Public Benchmarks

In this section, we evaluate Keye-VL-1.5 across various benchmarks. For general vision-language tasks, we select OpenCompass (Contributors (2023)), MMMU (Yue et al. (2024)), AI2D (Kembhavi et al. (2016)), MMBench (Liu et al. (2024b)), BLINK (Fu et al. (2024)), ZeroBench (Roberts et al. (2025)), VisuLogic (Xu et al. (2025b)), RealWorldQA (X (2025)), SimpleVQA (Cheng et al. (2025b)), MMStar (Chen et al. (2024d)), MMVP (Tong et al. (2024)), HallusionBench (Guan et al. (2024)) and OCRBench (Liu et al. (2024c)). For public Video tasks, we select Video-MME(Fu et al. (2025b)), Video-MMMU (Hu et al. (2025b)), TempCompass (Liu et al. (2024d)), LongVideoBench (Wu et al. (2024)), and MMVU (Zhao et al. (2025)). For MATH tasks, we select MathVision (Wang et al. (2024c)), MathVistaMINI (Lu et al. (2023)), MathVersevision (Zhang et al. (2024)), OlympiadBench (He et al. (2024)), WeMath (Qiao et al. (2024)), LogicVista (Xiao et al. (2024)), and DynaMath (Zou et al. (2024)).

We compare the performance of Keye-VL-1.5 in Thinking mode with Keye-VL-Preview and other state-of-the-art models of a similar scale, including Qwen2.5-VL 7B, InternVL3-8B (Zhu et al. (2025)), MiMo-VL-7B-RL 2508 (Xiaomi (2025)), and proprietary models such as GPT-4o and Claude-3.7-Sonnet.

On general vision-language tasks, Keye-VL-1.5 demonstrates competitive performance across most benchmarks, often achieving SOTA or near SOTA results and outperforming other models overall. On the large-scale general benchmarks OpenCompass, MMMUval\text{MMMU}_{\text{val}} and AI2D, Keye-VL-1.5 obtains scores of 79.5% 71.4% and 86.7% respectively, surpassing all other models. On MMBench and MMStar, Keye-VL also achieves the best performance. In mathematical reasoning tasks, Keye-VL-1.5 significantly outperforms Qwen2.5-VL 8B and InternVL3-8B, achieving comparable results with MiMo-VL 7B-RL.

In video-centric scenarios, Keye-VL-1.5 demonstrates superior capabilities compared to other open-source models. Our evaluations indicate that an accurate understanding of video content is Keye-VL-1.5’s core advantage. On public video benchmarks, Keye-VL-1.5 significantly outperforms other models, particularly on Video-MMMU, with an absolute improvement of 6.5%.

4 Internal Benchmarks

Despite extensive evaluations on a wide array of public video benchmarks, these benchmarks exhibit numerous limitations that necessitate a focused effort on developing a proprietary, internal evaluation suite. The primary issues are as follows:

Limited Task Coverage: Current publicly available benchmarks primarily focus on basic perception and simple reasoning capabilities, with insufficient coverage of specialized domains and temporal understanding tasks, failing to comprehensively evaluate model performance across diverse scenarios.

Oversimplified Question Formats: Existing evaluation tasks tend to employ overly simplistic questioning approaches. For instance, in video question answering tasks, queries often involve only the most basic inquiries about video content, such as counting the number of people present, which inadequately reflects real-world complexity.

Restrictive Answer Methodologies: To facilitate accuracy computation, questions are typically abstracted into yes/no responses or multiple-choice formats, which significantly deviate from natural user interaction patterns and limit the assessment of model’s genuine conversational capabilities.

Data Contamination Risks: Since datasets are publicly available, there exists a non-negligible possibility that models have already encountered these data during training, potentially leading to inflated performance metrics and compromised evaluation validity.

Language and Cultural Bias: Many existing test sets exhibit bias toward English-language scenarios, limiting our understanding of model performance in Chinese usage contexts and failing to capture culture-specific nuances and requirements.

Therefore, we construct a rigorous internal video evaluation benchmark. The video sources include both internal and external platform content, as well as artificially constructed videos, with resolutions ranging from 360p to 1440p, effectively avoiding overlap with existing training data. The questions are categorized into several dimensions to provide comprehensive coverage: Visual Element Recognition for assessing visual element identification capabilities, Reasoning Ability for evaluating logical reasoning skills, Temporal Info Understanding for measuring temporal information comprehension, Knowledge-based QA for testing knowledge-grounded question answering, Description Ability for evaluating descriptive capabilities, Robustness for testing model stability, Creative Ability for assessing creative thinking, and Domain Expertise for evaluating specialized domain knowledge.

The scoring methodology employs comparative evaluation across multiple model results and GSB (Good, Same, Bad) preference selection. The baseline models can be either GPT-4o or Gemini 1.5 Pro. The specific evaluation approach involves two methods. First, the scoring method uses multiple models (typically 2) to generate results that are evaluated separately on a 1-5 scale. Three annotators score the answers based on the video content and reference annotation guidelines, providing both fine-grained and overall scores. Second, the GSB method involves direct comparison between two model results using Good-Same-Bad preference selection. When two answers have significantly different scores, the higher-scoring answer is preferred. When the scores are similar, the selection is based on annotation rules and subjective judgment to determine which answer is better. If no clear distinction can be made, the selection reflects whether both answers are equally good, equally poor, or equally average based on answer quality.

5 Evaluation Results

Keye-VL-1.5-8B achieves significant performance improvements over previous versions: As demonstrated in Table 4, Keye-VL-1.5-8B establishes a substantial lead with an overall composite score of 3.53, representing a remarkable +0.51 improvement over Keye-VL-Preview. This advancement is particularly pronounced in correctness (+0.57) and completeness (+0.25), demonstrating the model’s enhanced ability to provide accurate and comprehensive responses. The model also shows notable gains in relevance (+0.11), indicating improved alignment between responses and user queries.

The model demonstrates competitive performance against industry benchmarks: In direct comparison with MiMoVL-7B-RL-2508, Keye-VL-1.5-8B achieves a higher overall score (3.53 vs. 3.40), establishing a +0.13 advantage in composite performance. The model particularly excels in correctness (+0.19) while maintaining competitive performance in completeness (-0.01). However, the evaluation reveals trade-offs in certain dimensions, with MiMoVL-7B-RL-2508 showing superior performance in fluency (+0.23), relevance (+0.08), and creativity (+0.15). This performance profile indicates that while our model achieves stronger factual accuracy, it faces challenges in language generation sophistication.

Detailed capability analysis reveals domain-specific strengths and optimization priorities: The fine-grained evaluation in Table 5 demonstrates Keye-VL-1.5-8B’s exceptional performance across multiple core capabilities. The model achieves decisive advantages in Reasoning Ability (3.81), Temporal Information Understanding (3.36), and Robustness (4.29), with the latter representing a substantial +0.83 lead over MiMoVL-7B-RL-2508. These results highlight the model’s particular strength in handling complex analytical tasks and maintaining consistent performance under challenging conditions. The model matches MiMoVL-7B-RL-2508 in Visual Element Recognition (3.49) and Creative Ability (3.66).

The model establishes a strong foundation in fundamental visual understanding capabilities: Keye-VL-1.5-8B’s performance demonstrates significant improvements in core visual processing tasks compared to previous iterations. The +0.35 advancement in visual element recognition and +1.00 improvement in reasoning ability over Keye-VL-Preview indicate substantial progress in fundamental perceptual and cognitive pathways. Particularly notable is the model’s +0.77 improvement in temporal information understanding, reflecting enhanced capability in processing sequential visual information and understanding dynamic relationships within video content. These foundational improvements provide a robust platform for handling complex multimodal reasoning tasks.

6 Ablation Studies and Findings

Table 6 presents a comprehensive evaluation of different training methodologies using varying quantities of high-quality data for SFT and MPO. The experimental results demonstrate that increasing the volume of SFT training data consistently enhances model performance across mathematical reasoning, logical inference, and OCR capabilities. Notably, our carefully curated preference dataset for MPO consistently yields additional performance improvements across all evaluated benchmarks. The implementation of Long CoT cold start training produces particularly remarkable results, with substantial performance gains observed across all benchmarks, most notably in mathematical reasoning tasks. These findings empirically validate the effectiveness of our proposed data processing pipeline and training methodology, demonstrating the synergistic benefits of combining high-quality supervised fine-tuning with preference optimization and strategic initialization approaches.

6.2 Effectiveness of Expert Models and Model Merging

Table 7 demonstrates the effectiveness of our expert model approach and model merging technique, using OCR tasks as a representative case study. Our base model initially achieved an average OCR performance of 78.25%, comparable to the preview version but exhibiting notable deficiencies in specialized domains such as license plate recognition, seal/stamp identification, and street scene text extraction. To address these limitations, we develop a specialized OCR expert model trained on curated domain-specific data. The OCR expert model demonstrates substantial improvements across all evaluated OCR benchmarks, achieving an average score of 83.65%. Furthermore, the strategic merging of our base model with the OCR expert yields additional performance enhancements, reaching an average score of 84.51%. This merged configuration significantly surpasses the perceptual capabilities of MiMo-VL, with particularly notable improvements in TextVQA (83.40% vs. 75.57%) and ChartQA (84.88% vs. 70.00%).

These empirical results validate the effectiveness of our proposed technical approach, demonstrating that domain-specific expert models can be successfully integrated with general-purpose base models to achieve superior performance across specialized tasks while maintaining overall model capabilities. Additionally, our experiments reveal the following findings:

Limited Training Steps: Expert models trained with more steps continue improving within their specialized domains. However, merged model performance initially increases with expert training steps, then decreases, indicating an optimal training duration.

Limited Learning Rate: Expert models achieve better performance with smaller learning rates, and the corresponding merged models also perform better.

The parameter divergence between expert and general models significantly affects merged model performance. Small divergences limit domain-specific improvements, while large divergences lead to suboptimal merged performance, creating a critical trade-off between specialization and integration.

6.3 Effectiveness of Alignment Reinforcement Learning

To validate the effectiveness of our alignment reinforcement learning approach, we conducted comprehensive evaluations starting from the Keye-VL-8B-preview baseline, focusing on instruction following capabilities and mathematical reasoning performance. Our evaluation framework encompasses both multimodal instruction following benchmarks (MIA-Bench and MMIFEval) and text-only instruction following assessments (IFEval and LiveBench). For mathematical reasoning evaluation, we selected four widely adopted benchmarks to ensure comprehensive coverage of mathematical capabilities. As demonstrated in Table 8, our alignment RL approach consistently outperforms the baseline across both inference modes. In the Think mode, substantial improvements are observed across all instruction following benchmarks, with notable gains of 4.35 points on MIA-Bench (91.95% vs. 87.60%), 6.48 points on MMIFEval (63.45% vs. 56.97%), and 5.40 points on LiveBench (64.70% vs. 59.30%). Similarly, in the No-Think mode, the model demonstrates consistent improvements, particularly achieving a 4.62-point enhancement on IFEval (78.37% vs. 73.75%). The mathematical reasoning capabilities also exhibit modest but consistent improvements across all evaluated benchmarks, with average gains ranging from 2-4 points. These results empirically validate that our alignment algorithm effectively enhances functional capabilities in instruction following while simultaneously strengthening general reasoning abilities. The consistent performance improvements across diverse evaluation metrics confirm the robustness and effectiveness of our alignment reinforcement learning methodology.

6.4 Effect of Partial Solutions During RL Phase

To evaluate the model’s performance under different hint conditions, the success rate of solving problems across four rollout attempts serves as the primary metric. Approximately 8,000 RL data samples are selected for testing, with the following conditions:

The model samples the same problem independently four times under each hint condition.

The score is determined by the number of correct answers (ranging from 0 to 4).

As shown in Table 9, without any hints, approximately 25.56% of the samples fail to provide a correct solution, significantly reducing the efficiency of the RL process. As the hints approach a complete solution (level 5), the error rate decreases, and the average score for the four attempts increases, indicating more stable and accurate responses. Additionally, a comparison between performance in the RL phase with and without partial solutions in Table 6 shows improvements across various benchmarks, including an increase in the average score from 79.41 to 80.13 on OpenCompass, and a 1.3-point improvement on MathVista, further validating the effect of partial solutions.

6.5 Impact of Rejection Sampling on SFT and RL Performance

In our RL iteration process, we employ rejection sampling twice. To validate the effectiveness of this approach, we conduct experiments starting with Keye-VL-8B-Preview, training it with the same RL dataset. In contrast, Keye-VL-8B-Preview-RFT-RL undergoes one round of iteration, followed by a second RL training phase. As shown in Figure 8, this iterative strategy significantly boosts RL performance, increasing the average mathematical benchmark score from 60.37 to 62.24, with similar improvements observed across general reasoning benchmarks. In Table 6, we compare the impact of various strategies, including Long CoT Cold Start, rejection sampling of SFT data using an RL model, and the subsequent selection of the best samples using a reward model for further SFT training (RFT-SFT). As a result, OpenCompass’s average score rises from 75.32 to 76.33, with consistent performance improvements across other benchmarks. Based on these findings, we adopt the SFT-RL-(RFT-SFT)-(RFT-RL) iterative model to further enhance performance.

Conclusion and Discussion

In this work, we presented Keye-VL-1.5, an advanced multimodal model that significantly enhances video understanding and vision-language tasks. By employing a novel Slow-Fast video encoding strategy, we efficiently balance temporal coverage and spatial resolution. The model’s progressive pre-training, with an extended context length, enables it to handle longer videos and complex visual content, while post-training methods focused on reasoning and human preference alignment improve instruction-following and reasoning abilities. Our evaluation demonstrates that Keye-VL-1.5 advances video understanding capabilities while maintaining strong performance on general vision-language tasks.

References

Appendix A Case Study

Appendix B Authors (Alphabetical order)

Core Contributors: Biao Yang, Bin Wen, Boyang Ding, Changyi Liu, Chenglong Chu, Chengru Song, Chongling Rao, Chuan Yi, Da Li, Dunju Zang, Fan Yang, Guorui Zhou, Guowang Zhang, Han Shen, Hao Peng, Haojie Ding, Hao Wang, Haonan Fan, Hengrui Ju, Jiaming Huang, Jiangxia Cao, Jiankang Chen, Jingyun Hua, Kaibing Chen, Kaiyu Jiang, Kaiyu Tang, Kun Gai, Muhao Wei, Qiang Wang, Ruitao Wang, Sen Na, Shengnan Zhang, Siyang Mao, Sui Huang, Tianke Zhang, Tingting Gao, Wei Chen, Wei Yuan, Xiangyu Wu, Xiao Hu, Xingyu Lu, Yi-Fan Zhang, Yiping Yang, Yulong Chen, Zeyi Lu, Zhenhua Wu, Zhixin Ling, Zhuoran Yang, Ziming Li.

Contributors: Di Xu, Haixuan Gao, Hang Li, Jing Wang, Lejian Ren, Qigen Hu, Qianqian Wang, Shiyao Wang, Xinchen Luo, Yan Li, Yuhang Hu, Zixing Zhang.