Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning
Maggie Huan, Yuetai Li, Tuney Zheng, Xiaoyu Xu, Seungone Kim, Minxin Du, Radha Poovendran, Graham Neubig, Xiang Yue
Introduction
Over the past years, the community has raced to push large language models (LLMs) to new heights on math-centric reasoning benchmarks such as MATH (Hendrycks et al., 2021b) and AIME. A steady stream of reasoning-tuned models (Muennighoff et al., 2025; Guha et al., 2025) now advances the state of the art on math benchmarks nearly every week, with some even surpassing the average performance of human experts (Team, 2025a; OpenAI, 2024).
The appealing performance on math reasoning is understandable: problems are well-posed, solutions are unambiguous, and evaluation is easily verifiable, often just a single number or expression (Luo et al., 2025). This clarity has made math a popular proxy task of LLM reasoning, and researchers have developed increasingly sophisticated training recipes to maximize model performance on mathematical reasoning (Wang et al., 2024a; Yue et al., 2024a; Luo et al., 2023; Shao et al., 2024; Wei et al., 2023).
This trend, on one hand, should be encouraged. Mathematics is often considered the foundational language of science, and enabling machines to reason precisely over math is central to the long-term vision of automated scientific discovery (Mishra et al., 2022). On the other hand, real-world tasks extend far beyond math. The majority of user-facing applications, question answering, dialogue, instruction following, require broader linguistic and commonsense competence that math alone does not test (Ma et al., 2025).
This raises a natural question: Do improved math reasoning abilities transfer to general LLM capabilities? Specifically, can gains in solving math problems transfer to other reasoning domains (e.g., scientific QA (Welbl et al., 2017), coding (Jain et al., 2025), agent planning (Xie et al., 2024), logical deduction (Dziri et al., 2024)) and to tasks (e.g., conversational QA (Reddy et al., 2019), instruction following (Zhou et al., 2023)) that do not require extensive reasoning?
To investigate, we evaluate over 20 representative open-weight reasoning models, all of which exhibit impressive performance on recent math benchmarks across a suite of other reasoning and non-reasoning tasks. To quantitatively analyze this problem, we propose Transferability Index, a metric to measure how reasoning models can transfer their capabilities from one domain to another. Surprisingly, as shown in Figure 2, we find that some of these models fail to transfer their improved mathematical reasoning capabilities to other domains while others succeed.
What drives this divergence? Model recipes vary widely in size, data distribution, and architecture. Yet among various parts, we identify one factor that consistently predicts transferability: the fine-tuning paradigm. Across families and sizes, models fine-tuned using reinforcement learning (RL) (Su et al., 2025; Yeo et al., 2025) exhibit much stronger generalization to non-math tasks than those trained with supervised fine-tuning (SFT) (Yue et al., 2024a, b), which often show signs of catastrophic forgetting over a wide range of non-math tasks.
To validate this observation, we conduct a controlled study. We fine-tune Qwen3-14B (Team, 2025b) on the high-quality math dataset derived from MATH and DeepScaler (Luo et al., 2025). For SFT, we construct targets via rejection sampling using Qwen3-32B, keeping only teacher responses that yield correct final answers. For RL, we apply a standard GRPO (Shao et al., 2024) setup using answer correctness as the reward. As shown in Figure 1, the results mirror our large-scale audit: RL-tuned models generalize well to non-math domains, despite being trained solely on math queries, while SFT-tuned models do not.
To better understand why this occurs, we use two diagnostic tools: (1) latent-space principal component analysis (PCA) on hidden states across layers, and (2) KL-divergence on token distributions before and after fine-tuning. These methods allow us to quantify how much the model’s internal representations and output space shift during training. We find that SFT induces significant drift in both latent and output spaces, especially for non-reasoning inputs, whereas RL better preserves the geometry of internal features and the stability of the token distributions.
Phenomena: Performance Discrepancies of Reasoning Models
Setup. We evaluate over 20 off-the-shelf reasoning models on different benchmarks. Namely, we select benchmarks from the following three categories: (1) math reasoning tasks: MATH500 (Hendrycks et al., 2021b), AIME24, AIME25, OlympiadBench (He et al., 2024), which contain mathematical problems only; (2) other reasoning tasks: LiveCodeBench (Jain et al., 2025), GPQA-Diamond (Rein et al., 2024), ACPBench (Kokel et al., 2025), HeadQA (Vilares and Gómez-Rodr´ıguez, 2019), which contain more general reasoning questions, such as medical reasoning, code generation, and language-based agent planning tasks; (3) non-reasoning tasks: CoQA (Reddy et al., 2019), IFEval (Zhou et al., 2023), HaluEval (Li et al., 2023), MC-TACO (Zhou et al., 2019), which contain factual, alignment, or conversational problems such as commonsense question answering and instruction-following. We used accuracy to evaluate the models’ performance. Detailed explanation about experiment setup, benchmarks, and evaluation metrics can be found in Appendix A.2.
To better evaluate the model’s transferability across a wide range of task groups, we define Transferability Index (TI) as follows:
Next, the two Transferability Indices are
The TI value is compared against 0, any TI above zero indicates positive transfer observed.
2 Control Study
Motivated by our findings in Section 2.1, we design a light-weight controlled study to directly compare SFT and RL on an identical dataset. Concretely, we start from a small, high-quality mathematics dataset (see Appendix A.2.2 for details), then query a strong teacher model (Qwen3-32B-Instruct) to extract complete chain-of-thought (CoT) reasoning traces with reject sampling. These CoT traces become our SFT training targets, while the original answer labels serve as the rewards for RL. This alignment ensures both paradigms learn from the same data samples. Then, we take the Qwen3-14B-Base model and fine-tune it in two ways: (i) SFT on the teacher-generated CoT traces; (ii) RL using only the groundtruth. We name our model UniReason. We compare against the Qwen3-14B-Base model. Evaluation is conducted on three benchmark groups mentioned above using accuracy. Details about training datasets, baseline models, and hyperparameters could also be found in Appendix A.2.
Our experimental results on three groups of benchmarks (see Table 1) reveal a consistent pattern:
On math reasoning (Table 1), our UniReason-Qwen3-14B(RL) model climbs to 55.7% on AIME24, 87.8% on MATH500, and 33.8% on OlympiadBench, outperforming corresponding SFT-based models.
For other reasoning tasks, SFT-based models make uneven progress (e.g. UniReason-Qwen3-14B-SFT-think scores 55.9% on GPQA), whereas RL fine-tuning yields significant lifts: UniReason-Qwen3-14B(RL) gains 1.8% on GPQA, and 17.1% on LiveCodeBench2 over SFT.
Crucially, in non-reasoning evaluations, SFT models stagnate or decline, while the RL model recover and exceed the base in nearly all the benchmarks.
These results show that RL-tuned reasoning models perform generally better than SFT-based models on both reasoning and non-reasoning tasks when carefully controlling other factors. Especially, our UniReason model is trained on a single distilled math dataset, but it still preserves and even improves general-domain performance while showing strong reasoning gains.
Latent Representation Shifts: Insights from PCA Analysis
As discussed in Section 2.1, applying SFT to the Qwen model improves reasoning abilities such as mathematical problem-solving and code generation, but substantially impairs general-domain performance. We observe that most SFT models fail to transfer their improved mathematical reasoning capabilities to other domains. In contrast, our controlled study shows that RL-tuned models generalize well to non-math domains, despite being trained solely on math queries, whereas SFT-tuned models do not.
To understand the underlying cause of this transferability gap, we employ PCA shift analysis to examine how the internal feature geometry of the model evolves under different training paradigms, model sizes, and model families across diverse query distributions. Recent studies (Xu et al., 2025a; Zheng et al., 2025) demonstrate that PCA shift analysis provides a sensitive and interpretable measure of representational changes relevant to task performance. Importantly, changes in model parameters do not always correspond to functional differences: large weight updates may leave outputs unchanged, while subtle parameter modifications can lead to significant shifts in the activation distribution. By focusing on hidden representations, PCA shift directly captures how the model encodes and processes information, offering a more faithful account of its internal knowledge state than parameter-based metrics. This perspective allows us to distinguish between true knowledge erasure and parameterization changes that leave the underlying feature space intact. Furthermore, since transferability fundamentally relies on the alignment and stability of learned representations across tasks or domains, PCA shift is particularly effective for diagnosing changes that may impact cross-domain generalization. Shifts in principal components reveal whether the model’s internal feature space remains suitable for knowledge transfer or has been disrupted by training or unlearning interventions.
Models and Tasks. In Section 2.1, we observe that models trained on math datasets show moderate transferability on other reasoning tasks. We perform PCA shift analyses on the corresponding models and tasks, aiming to critically assess the robustness of these phenomena from a feature-space perspective.
Evaluation. Given input queries , we extract hidden states at each layer for each model state . Applying PCA () to , we compute the mean projection onto the first principal direction (PC1) and onto the second (PC2). The PCA shift is defined as for PC1, while for PC2, we directly report as an auxiliary indicator of distributional change. Small shifts indicate stable features.
2 Investigating Latent Space Shift
To quantify the overall latent shift, we define a representation center for each model state as the mean of PCA-projected coordinates across all layers: where denotes the total number of layers and is the vector of PCA coordinates for layer in state . The latent shift between two model states, such as the original (base) and an updated model, is then measured by the Euclidean distance:
Based on the analyses in Appendix A.3, RL-based training proves essential for developing robust and generalizable language models that maintain a strong balance between general-domain and reasoning capabilities. Motivated by this, we further analyze our finetuned models proposed in the controlled study. As shown in Table 2, RL-based models (highlighted in red) achieve the lowest PCA shift magnitudes across math, other-reasoning, and non-reasoning tasks. Figure 3 further supports these findings, illustrating that the RL-based model consistently yields minimal and tightly clustered latent shifts across diverse benchmarks. In contrast, SFT-based model, particularly those without explicit reasoning signals, exhibit more scattered and pronounced shifts.
These results, along with the evaluations in Section 2.1, highlight the clear advantage of RL compared with SFT. This confirms the necessity of a holistic and well-balanced optimization objective, rather than isolated interventions, to minimize catastrophic forgetting and preserve performance in large-scale language models.
Token Distribution Shifts: Insights from KL Divergence and Rank Analyses
In this section, we conduct token-level analyses to further examine the distribution shift of RL and SFT models trained on mathematical reasoning data.
KL-divergence serves as a standard metric for measuring differences between probability distributions. For token rank shift analysis, we first generate tokens using the fine-tuned model, then decode these same tokens using the backbone model to determine their original ranking positions. The rank shift is calculated as the difference in token rankings between the fine-tuned model and the backbone model for each token (Li et al., 2025c; Lin et al., 2023). Following the observations in Section 2.1, we perform additional token-distribution analyses on the corresponding models and tasks to assess the model distribution shift from a token-space perspective. Specifically, we employ KL-divergence and token rank shift metrics to analyze distribution shifts between models.
RL models exhibit lower KL-divergence from backbone models. In Figure 5, we observe that the KL divergence of SFT models on both reasoning and non-reasoning tasks is significantly larger than that of RL models. This indicates that RL models exhibit substantially less distribution shift from the token distribution level compared to SFT models. For instance, UniReason-Qwen3-14B-SFT-no-think demonstrates KL divergences of 0.372 and 0.283 on MATH-500 and IFEval respectively compared to the backbone model, whereas UniReason-Qwen3-14B(RL) achieves considerably lower KL divergences of only 0.084 and 0.019 on the corresponding tasks.
RL models demonstrate reduced token rank shifts. In Figure 15, we further analyze the token rank shift for both SFT and RL models. We find that RL models exhibit substantially lower average token rank shifts compared to the backbone model than SFT models. Specifically, UniReason-Qwen3-14B(RL) demonstrates an average token rank shift of only 0.98, while the UniReason-Qwen3-14B-SFT-no-think shows a dramatically higher average token rank shift of 10.6. This suggests that SFT models experience greater token distribution shifts than RL models across both reasoning and non-reasoning tasks.
Figure 6 provides a detailed visualization of token rank shifts across different position indices for both reasoning and non-reasoning tasks. We observe that RL models exhibit less token rank shifts (less than 10) at only a few positions. In contrast, SFT models demonstrate substantial rank shifts across numerous positions throughout the sequence.
RL models selectively shift task-relevant tokens, while SFT models shift numerous irrelevant tokens. Table 4.1 presents a comprehensive case study examining shifted tokens in RL and SFT models for reasoning and non-reasoning queries. For RL models, we observe highly selective token shifting, with only a small number of query-relevant tokens undergoing shifts. In reasoning queries, shifts are limited to essential logical tokens such as "define", "add", "second", and "number" while in non-reasoning queries, only task-specific keywords like "<<" ">>" "write" and "formally" experience rank changes. In contrast, SFT models exhibit extensive token shifts, with 390 and 158 token shifts in reasoning and non-reasoning queries respectively. These shifts include numerous query-irrelevant tokens. For example, non-reasoning queries inappropriately introduce reasoning tokens, leading to unnecessary overthinking that detracts from performance. Please see Appendix A.4 for the full model response. Also, we calculate the word frequency of shifted tokens from both RL-tuned and SFT-tuned models under math reasoning task, merging them into one pool and select top 250 frequent tokens in it. Then we plot a word cloud for the selected tokens, as shown in Figure 4. The figure confirms our observation that RL selectively shifts task-related tokens while SFT shows both relevant and irrelevant token shifts.
Reasoning Fine-Tuning of LLMs. Recent advancements in large language models have notably emphasized specialized fine-tuning methods to enhance reasoning capabilities (Wong et al., 2025; Chen et al., 2023; Ziegler et al., 2019; Liu et al., 2025; Wang et al., 2024c; Li et al., 2025b; Feng et al., 2025; Xu et al., 2025b; Yang et al., 2025; Li et al., 2025a; Yeo et al., 2025). The chain-of-thought prompting strategy introduced by Wei et al. (2022) encourages models to produce step-by-step explanations, significantly boosting performance in symbolic reasoning tasks (Lambert et al., 2025; Wei et al., 2022; Longpre et al., 2023; Yu et al., 2024). Subsequent extensions, such as DeepSeek-R1 (Team, 2025a), have integrated reinforcement learning approaches alongside CoT, optimizing models through reward-driven policy improvements. Such RL-enhanced fine-tuning has achieved state-of-the-art results on benchmarks and competitive programming challenges (Hendrycks et al., 2021b; Team et al., 2025; Team, 2025a; Lambert et al., 2025).
Supervised Fine-Tuning vs. Reinforcement Learning for LLMs. Fine-tuning methods for reasoning typically fall into two major categories: supervised fine-tuning and reinforcement learning (Chen et al., 2024). SFT methods predominantly utilize annotated reasoning trajectories or solution traces, directly training models to replicate explicit reasoning sequences from datasets (Wei et al., 2022; Wang et al., 2023). RL-based fine-tuning, however, guides models by rewarding accurate and logically coherent reasoning steps without explicit step-by-step supervision, allowing exploration and optimization of reasoning pathways through feedback loops (Ziegler et al., 2019; Liu et al., 2025; Wang et al., 2024c; Chu et al., 2025).
Generalization in Reasoning Models. Interestingly, models heavily fine-tuned for formal reasoning sometimes falter on more general language tasks (Kumar et al., 2022). For example, OpenAI’s o1, while excelling in STEM benchmarks, raised concerns about its versatility on other tasks (OpenAI, 2024). Follow-up research introduced reinforcement fine-tuning precisely to address this gap, aiming to adapt a generalist model’s reasoning to new domains with limited data (Zhang et al., 2024). Indeed, o1 and similar reasoning models are built on strong general-purpose bases to retain broad knowledge (OpenAI, 2024; Hendrycks et al., 2021a). Nonetheless, trade-offs have been observed. Wang et al. (2024b) found that fine-tuning on a narrow set of instruction types can degrade a model’s performance on other skills. Recent works have also stepped into analyzing the cross-domain performance of reasoning models (Sun et al., 2025), especially for RL-based approaches (Cheng et al., 2025; Hu et al., 2025).
Representation-Level Analysis. Fine-tuning for reasoning models not only boosts task performance but also alters the model’s internal representations (Sheng et al., 2024). Recent studies have begun to probe how CoT-based fine-tuning changes the latent space of LLMs (Xu et al., 2025a; Wang et al., 2025). Lobo et al. (2024) find that task-specific fine-tuning can reduce the faithfulness of a model’s generated reasoning chains, indicating shifts in its underlying inference mechanisms. Complementary analyses of hidden states provide insight into such shifts. Xu et al. (2024) proposes a quantitative framework for assessing ideas that leverages hidden representations from LLMs to predict the merit of scientific ideas. Techniques like principal component analysis further reveal that fine-tuning can carve out new directions in representation space that correspond to reasoning-related features (Xu et al., 2025a; Zhou et al., 2025).
In this work, we investigated the factors that influence the transferability of reasoning models across reasoning and non-reasoning benchmarks. Our key findings are as follows. First, besides model size and model architecture, the choice of fine-tuning paradigm strongly shapes transfer: RL-tuned models achieve significant gains on math reasoning while preserving positive transfer to other reasoning tasks and non-reasoning tasks, whereas SFT often incurs negative transfer on non-reasoning benchmarks. Second, PCA analysis of latent space confirms that RL induces minimal drift from backbone representations thus maintaining feature stability, while SFT produces larger latent shifts, especially in non-reasoning domains. Third, token-distribution analysis shows that RL selectively adjusts only a handful of task-relevant tokens, whereas SFT perturbs many irrelevant tokens, indicating RL’s more targeted optimization. Notably, our UniReason-Qwen3-14B (RL) fine-tuned on 47K math examples achieves the best balance of reasoning improvement and general-domain retention among compared models, strongly validating our hypotheses and analysis.
Maggie Huan co-led the project, co-prepared the dataset, led initial experiments and analyses, conducted benchmark evaluations, wrote Sections 2, 5, and 6 and revised the paper.
Yuetai Li co-led the project, conducted the preliminary study and initial testing, initially identified the phenomenon (Section 2), analyzed results, performed the token-level analysis (Section 4), and contributed to method refinement.
Tuney Zheng co-led the project, co-prepared the math dataset, rl models and fine-tuned models for the controlled study, contributed to result analysis, and supported method development.
Xiaoyu Xu performed latent-space PCA analysis, contributed to early result analysis, and wrote Section 3.
Seungone Kim, Minxin Du, and Radha Poovendran provided valuable feedback on the manuscript.
Graham Neubig co-supervised the project and offered substantive guidance and feedback throughout.
Xiang Yue conceived and co-led the project by proposing the initial research ideas, designing the methodology, analyzing results, writing Abstract and Introduction, revising the paper, and supervising the overall work.
The authors would like to thank Qian Liu for helping set up the training environment. This work was supported in part by a Carnegie Bosch Institute Fellowship to Xiang Yue.
Appendix A Appendix
As discussed in Section 2.1, we provided the complete evaluation for the Transferability Index for the off-the-shelf models on other reasoning and non-reasoning tasks in Table 4.
A.2 Full Evaluation Setup
Reinforcement Learning has recently proven effective at steering large language models toward complex, multi-step objectives by optimizing policies with scalar reward signals (Zeng et al., 2025). We used the Verl framework(Sheng et al., 2025) and GRPO (Shao et al., 2024) setup to fine-tune the Qwen-3-14B-Base model, utilizing answer correctness as the reward signal. Our RL uses a learning rate of with an overall train batch size of 512 and clipping thresholds set between 0.22 and 0.28. We generate sequences up to 16k tokens long and perform 16 rollouts per prompt, then update the model in mini-batches of 128 samples. Both KL-divergence and entropy penalties are turned off (coefficients set to zero). We train the model for 140 steps and used the corresponding checkpoint.
Supervised Fine-Tuning remains a fundamental technique for adapting large pre-trained models by directly minimizing cross-entropy on high-quality datasets (Chu et al., 2025). We use the LLaMA-Factory framework (Zheng et al., 2024), which is an extensible and user-friendly framework supporting multiple architectures and advanced optimization algorithms, to fine-tune our model on teacher-generated chain-of-thought traces. We use as learning rate, the batch size is 512 and we train for 1.5 epoch to align with our RL settings.
A.2.2 Training Datasets
As briefed in Section 2.2, our base training dataset is a curated set of 47K high-quality mathematics problems. We stratified the examples using two complementary sources: low-difficulty problems drawn from the DeepScaler dataset (Luo et al., 2025), and high-difficulty (levels 3–5) problems extracted from SimpleRL (Zeng et al., 2025). To generate CoT annotations, we prompt each problem into the Qwen3-32B-Instruct model (Team, 2025b) and use reject sampling to generate our dataset.
To further explore the effect of training data distribution for SFT-based reasoning models, we also distill a larger and more comprehensive dataset collected from General-Reasoner (Ma et al., 2025), which contains 232K examples across reasoning and non-reasoning tasks (e.g., Math, Chemistry, Business). This additional distilled set is used to train the General-Reasoner model using supervised fine-tuning.
A.2.3 Baselines
In our experiments, we compare against Qwen3-14B-Base model (Team, 2025b), which is the original Qwen3-14B model without any further adaptation. This serves as the unmodified backbone for all fine-tuning models. Also, we report the results of Qwen3-14B-Instruct model under our tested benchmarks. It is an instruction-tuned version of Qwen3-14B model trained on a large, general-purpose instruction-following dataset. We evaluate it under two prompting modes:
think: prompts include a special
no-think: prompts are provided without the
Due to the wide range and enormous training data, this model is considered to give the optimal outputs across tasks in the current 14B-series models.
To further validate our observation, we apply our controlled study pipeline also for General-Reasoner (Ma et al., 2025), an RL-tuned reasoning model that are also math experts. We distilled the dataset for SFT finetuning with the same rejection sampling method using their proposed dataset. The dataset contains 232K samples covering various reasoning and non-reasoning tasks. Then we finetune the Qwen3-14B base model using the distilled dataset and name the model as General-Reasoner-Qwen3-14B(SFT) to directly compare with the RL-based General-Reasoner for a fairer and more comprehensive controlled study towards SFT and RL. The Transferability Index results could be retrieved in Table 4, the obeservation also confirms our initial hypthotesis.
A.2.4 Evaluation Benchmarks
In the experiment, we evaluated our model across a wide range of benchmarks. Notably, to explicitly reveal the transferability of reasoning models, we grouped them into three categories by their content:
We collected the following datasets that are composed of mathematical problems, which means that they typically need a mathematical reasoning process to get the answer:
MATH500 (Hendrycks et al., 2021b): A curated subset of 500 problems sampled from the broader MATH dataset, covering topics like algebra, combinatorics, geometry, and number theory.
AIME: Problems drawn from the American Invitational Mathematics Examination (AIME) 2024 and 2025, each with 30 challenging short-answer questions requiring multi-step reasoning.
OlympiadBench (He et al., 2024): Problems sourced from international olympiads (e.g., IMO and regional contests).
We collected the following datasets that are mainly composed of general reasoning problems containing a wider range of subjects:
LiveCodeBench (Jain et al., 2025): It is a continuously updated, contamination-free coding benchmark. We used its second version.
GPQA-Diamond(Rein et al., 2024): It is a graduate-level question-answering dataset that contains multiple-choice questions in biology, physics, and chemistry. We followed its diamond split.
ACPBench (Kokel et al., 2025): It has 7 atomic reasoning tasks around 13 classical planning domains. We only used the multiple-choice problems.
HeadQA (Vilares and Gómez-Rodr´ıguez, 2019): Multiple-choice QA from healthcare-specialist certification exams, including questions across pharmacology, chemistry, nursery, psychology, biology, and medicine.
We collected the following datasets that are mainly composed of problems with factual answers, which means that they do not need a reasoning process to give the answer:
CoQA(Reddy et al., 2019): It has 127K questions in dialogues over passages, focusing on maintaining context and coreference across turns.
IFEval (Zhou et al., 2023): It contains over 500 prompts, each embedding verifiable instructions. Evaluates strict vs. loose adherence to instructions.
HaluEval (Li et al., 2023): It contains human-annotated samples where models must distinguish factual content from hallucinations.
MC-TACO (Zhou et al., 2019): It is a multiple-choice benchmark designed to evaluate models’ temporal commonsense, covering duration, ordering, typical time, frequency, and stationarity.
A.2.5 Evaluation metrics
We used LLM-Harness (Gao et al., 2024) to evaluate the models’ performance on OlympiadBench, ACPBench, HeadQA, CoQA, HaluEval, MC-TACO and used Eval-Chemy (Raoof et al., 2025) MATH500, AIME24, AIME25, GPQA-Diamond, LiveCodeBench, IFEval. On MATH500, AIME24, AIME25, GPQA-Diamond, and LiveCodeBench, we used 0.6 as temperature, and 0.95 as top-p value. In our experiments, we used accuracy to evaluate the models’ performance. Specifically, for AIME24 and AIME 25, we averaged accuracy on 10 samples. For GPQA-Diamond, LiveCodeBench and MATH 500, our score is the average accuracy over 3 samples. Specifically, we used version 2 and overall accuracy on LiveCodeBench. For ACPBench, we only used multiple choices, and averaged the score for all 10 tasks as the final score. For OlympiadBench, we only used math queries in English, and thus categorized Olympiad as a math benchmark. For HaluEval, the performance is the accuracy averaged on 3 tasks with zero-shot. And for IFEval, we used strict instruction accuracy as the score. For OlympiadBench, ACPBench, HeadQA, CoQA, HaluEval, and IFEval, we used greedy sampling and sampled only once.
A.3 PCA Analysis under Varying Settings
Table 5 summarizes the mean across math, other-reasoning, and non-reasoning tasks, providing an overall assessment of latent-space shifts under different training paradigms. Figure 7, 8, 9, 10 illustrates the paradigm comparison between SFT and RL, Figure 11 contrasts model sizes (7B versus 32B), and Figure 12 compares model families (Qwen vs. Llama).
Impact of Training Paradigm. RL-based fine-tuning consistently results in lower PCA shifts than SFT across math, other-reasoning, and non-reasoning tasks. As shown in Table 5, models such as SimpleRL-7B and SimpleRL-14B exhibit significantly smaller feature shifts compared to their SFT-trained counterparts. This observation is further visualized in Figure 7, 8, 9, 10, which demonstrates more concentrated representation shifts under RL. These findings are consistent with the phenomena discussed in Section 2.1, reinforcing that RL enhances generalization by better preserving internal representations. Overall, these results suggest that RL is substantially more effective at preserving general-domain representations and mitigating catastrophic forgetting.
Effect of Model Size. Increasing model size from 7B to 32B substantially reduces PCA shifts on math, other-reasoning, and non-reasoning tasks, as shown in Table 5. Figure 11 further confirms that larger models exhibit more concentrated and stable latent representations in non-reasoning domains, indicating reduced feature drift. These findings suggest that scaling up model size is particularly effective for enhancing stability and resistance to latent space drift in general-domain tasks.
Effect of Model Family. The results in Table 5 and Figure 12 indicate that both RL-tuned Llama and Qwen models exhibit small PCA shifts. While Qwen displays similar shifts as Llama on reasoning tasks, it also has higher drift on other-reasoning and non-reasoning tasks. In contrast, Llama models show greater resilience to drift in general domains. These differences stem from the distinct architectures and pre-training data, suggesting that RL strategies should be tailored to each model’s characteristics.
Summary. Our analyses show that latent space drift is mainly determined by training paradigm, model size, and model family. RL-based fine-tuning is the most effective at preserving general-domain representations and mitigating catastrophic forgetting. While larger models offer greater feature stability, especially on other-reasoning and non-reasoning tasks, scaling alone is insufficient. Qwen models remain sensitive to reasoning-related pre-training data, underscoring the need for tailored strategies. Overall, RL-based training is critical for developing robust language models that balance general-domain coverage with reasoning ability.
A.4 Token Distribution Shifts
This part provides the complete queries and responses in our case study in Section 4.
Query: Solve the following math problem. Present the final answer in the format: Final Answer: boxed{your_answer}. Problem: Ten treeks weigh as much as three squigs and one goolee. Two treeks and one goolee are equal in weight to one squig. The combined weight of how many treeks equals the weight of one squig?
RL-tuned Models: Let’s define the variables for the weights of the treeks, squigs, and goolees. Let: - be the weight of one treek, - be the weight of one squig, - be the weight of one goolee. We are given two pieces of information: 1. Ten treeks weigh as much as three squigs and one goolee: 2. Two treeks and one goolee are equal in weight to one squig: We need to find how many treeks combined equal the weight of one squig. In other words, we need to find a number such that . First, let’s solve the second equation for : Now, substitute into the first equation: Now, add to both sides to isolate the terms involving : Now, divide both sides by 4 to solve for : This tells us that the weight of one squig is equal to the weight of three treeks. Therefore, the combined weight of three treeks equals the weight of one squig. Final Answer:
SFT-tuned Models: Okay, let’s see. I need to solve this problem where they talk about treeks, squigs, and goolees. The question is asking how many treeks combined equal the weight of one squig. First, let me parse the problem again. The first statement says ten treeks weigh as much as three squigs and one goolee. So, , where is treek, is squig, is goolee. The second statement is two treeks and one goolee equal one squig. So, . The question is asking for how many treeks equal one squig. So, find such that . Alright, so I have two equations: 1) 2) I need to solve these equations to find where . Let me see. Maybe I can express from the second equation and substitute into the first. From equation 2: . Let me check that. If , then subtract from both sides: . Yes, that’s right. Now substitute this into equation 1. So equation 1 becomes: Simplify the right side: So now the equation is Let me add to both sides: Divide both sides by 4: . Wait, so ? Wait, hold on. The question is asking for how many treeks equal one squig. If , then 1 squig equals 3 treeks. Where did I go wrong? Let me check my equations again. The problem says: First statement: Ten treeks weigh as much as three squigs and one goolee. So . Correct. Second statement: Two treeks and one goolee are equal in weight to one squig. So . Correct. Then solving for from the second equation: . Substitute into the first equation: . So . Therefore, 1 squig equals 3 treeks. Maybe I misread the problem? Wait, maybe the problem is asking for how many treeks equal one squig, but in the equations, maybe I have to express in terms of . Let me try solving for . From , then . So one treek is of a squig. Therefore, to get one squig, you need 3 treeks. So answer is 3. Am I wrong? Rewrite equations: Equation 1: Equation 2: Substitute into equation 1: Still the same. Try elimination form: Equation 1: Equation 2: Add equations: Still the same. Try plugging in values: assume , then , and from equation 2: Check equation 1: , consistent. Conclusion: equations are consistent, and answer is .
Query: Write an email to my boss telling him that I am quitting. The email must contain a title wrapped in double angular brackets, i.e. <
RL-tuned Models: Write an email to my boss telling him that I am quitting. The email must contain a title wrapped in double angular brackets, i.e. <
SFT-tuned Models: Okay, so I need to write an email to my boss telling him I’m quitting. The email has to have a title wrapped in double angular brackets like <
A.5 Breakdown Evaluation Results
As discussed in the paper, we evaluated recent off-the-shelf models and our fine-tuned models across math reasoning, other reasoning and non-reasoning tasks. The detailed breakdown results are presented in Table 6, 7 and 8: