Generative Verifiers: Reward Modeling as Next-Token Prediction

Lunjun Zhang, Arian Hosseini, Hritik Bansal, Mehran Kazemi, Aviral Kumar, Rishabh Agarwal

Introduction

While large language models (LLMs) demonstrate remarkable capabilities, they often confidently make logical and factual mistakes (Zhang et al., 2023). These mistakes pose a significant challenge for reasoning problems, where a single mistake can invalidate the entire solution. A common strategy to address this issue is Best-of-N (Charniak and Johnson, 2005; Cobbe et al., 2021): the LLM generates N candidate solutions for a given problem, and a learned reward model, referred to as a “verifier”, ranks these solutions and picks the most suitable one. The effectiveness of this strategy hinges on how accurate the verifier is, making it crucial to identify better approaches for training verifiers.

On reasoning domains, LLM-based verifiers are typically trained as discriminative reward models (RMs) to assign numerical scores to candidate solutions, which is then used to classify them as correct or incorrect (Cobbe et al., 2021; Lightman et al., 2023; Wang et al., 2023). However, this scoring approach does not utilize the text-generation capabilities LLMs are fundamentally designed for. As a result, discriminative RMs miss out on the inherent strengths of generative LLMs, such as unified instruction tuning (Chung et al., 2022), chain-of-thought reasoning (Wei et al., 2022), and utilizing additional inference-time computation for better performance (Wang et al., 2022; Brown et al., 2024). While LLM-as-a-Judge (Zheng et al., 2024), which simply prompts off-the-shelf generative LLMs, also offer the above advantages, it often underperforms trained LLMs-based verifiers, especially on reasoning.

In this work, we propose training verifiers with next-token prediction, which we call GenRM, to leverage the text generation capabilities of LLMs (Figure 2). Concretely, to produce a numerical score for a solution, the verifier now uses a prompt such as ‘Is the answer correct?’, and represents the score as the probability of a single text token (e.g., ‘Yes’ or ‘No’) under the context and the prompt. GenRM naturally supports CoT reasoning (Wei et al., 2022): it can be trained to reason explicitly by generating a verbalized rationale before predicting correctness using ‘Yes’ or ‘No’ token (Figure 3), assuming rationales are available during training. We can further boost verification accuracy of CoT verifiers using majority voting (Wang et al., 2022): sampling multiple CoT rationales (votes) and calculating the average score of the ‘Yes’ token across all samples, resulting in a favorable use of more test-time computation. Moreover, GenRM’s next-token prediction training enables unifying solution generation with verification, which has been difficult (Hosseini et al., 2024), potentially improving verification through positive knowledge transfer from solution generation.

GenRM outperforms discriminative RMs, LLM-as-a-Judge, and self-consistency on algorithmic string manipulation and math reasoning tasks (Figure 1), namely Last Letter Concat (Wei et al., 2022), Word Sorting from Big-Bench Hard (Suzgun et al., 2022), and GSM8K (Cobbe et al., 2021). Best-of-N performance with GenRM-CoT further improves when using majority-voting, nearly matching performance with oracle verifier on algorithmic reasoning tasks. On GSM8K (Figure 1, right), when using a Gemma-9B GenRM-CoT model to verify the outputs of Gemini 1.0 Pro, we observe a 20% improvement in terms of the number of problems solved (73%→92.8%73\%\rightarrow 92.8\%), surpassing GPT-4 and Gemini 1.5 Pro. Moreover, we find that generative verifiers exhibit scales favorably as we increase dataset size as well as model capacity. Furthermore, GenRM-CoT also outperforms LLM-as-a-Judge as we scale inference-time compute by sampling multiple verification rationales for majority voting. Overall, these results suggest that generative verifiers hold significant potential for improving the reasoning capabilities of LLMs.

Preliminaries

An autoregressive language model generates an output sequence y=(y1,y2,…,yT){\mathbf{y}}=(y_{1},y_{2},\ldots,y_{T}) given a input context x{\mathbf{x}} (e.g., math problem) by predicting tokens one at a time, based on the previously generated tokens. Assuming that the language model is parameterized by θ\theta, the conditional probability distribution of generating a sequence y{\mathbf{y}} given context x{\mathbf{x}} is

Next-token prediction is the typical approach for pre-training and fine-tuning LLMs. In particular, supervised fine-tuning (SFT) minimizes the cross-entropy loss between the model’s predicted next token and the actual target token in a given sequence. Given a dataset D={(x,y)}{\mathcal{D}}=\{(x,y)\} of input context x{\mathbf{x}} and target response y{\mathbf{y}}, the SFT loss is given by:

Best-of-N is a widely-used approach to improve the reasoning performance of LLMs (Cobbe et al., 2021; Lightman et al., 2023). Specifically, given a test problem, we sample N candidate solutions from a generator LLM. These candidates are then scored using a learned verifier or reward model, and the highest-scoring solution is selected as the final answer. A better verifier increases the chance of selecting the correct solution, improving test accuracy.

Discriminative Verifiers. The prevalent approach of training verifiers for reasoning domains is to fine-tune an LLM as a classifier on a dataset of correct and incorrect solutions generated from a fixed LLM, using the binary cross-entropy loss. To do so, these verifiers directly assign a numerical score rθ(x,y)∈r_{\theta}({\mathbf{x}},{\mathbf{y}})\in to estimate the probability that a solution y{\mathbf{y}} is correct for a problem x{\mathbf{x}}. As such, these verifiers do not utilize the text generation the capabilities of LLMs. Given a reward-modeling (RM) dataset {\mathcal{D}}_{RM}={\mathcal{D}}_{\text{incorrect}}\mathbin{\scalebox{1.1}{\bigcup}}{\mathcal{D}}_{\text{correct}}, we train discriminative RMs as follows:

where y+{\mathbf{y}}^{+} are correct and y−{\mathbf{y}}^{-} are incorrect solutions, and clscls corresponds to a special vocabulary token. In this work, we always use a balanced data mixture between correct (Dcorrect{\mathcal{D}}_{\text{correct}}) and incorrect (Dincorrect{\mathcal{D}}_{\text{incorrect}}) problem-solution pairs.

GenRM: Verification as Next-Token Prediction

Discriminative LLM-based verifiers (3) do not utilize the text generation capabilities of pretrained LLMs. To address this issue, we propose training verifiers or that can generate text, which we call GenRM, using standard next-token prediction (2). To do so, GenRM represents solution correctness using the LLM’s probability distribution over tokens, instead of predicting a separate numerical score. This keeps the generation abilities of GenRM intact as the verification decision is just another token, while also enabling several advantages that come for “free” with LLMs such as unified training for solution generation and verification, chain-of-thought reasoning, and inference-time computation.

At inference, we use the likelihood of the ‘Yes’ token as the verifier’s score for re-ranking solutions:

This score takes into account the verifier’s confidence about its correctness prediction, which reduces the chance of being miscalibrated and wrong at test-time when using a binary ‘Yes’ or ‘No’ prediction.

2 Unifying Generation and Verification

GenRM seamlessly integrates reward modelling, which distinguishes between correct and incorrect solutions, with SFT for generating correct solutions. This can be done by simply changing the data mixture in the SFT loss (2) to include both verification and generation tasks. Given a verification dataset Dverify{\mathcal{D}}_{\text{verify}}, which can be DDirect{\mathcal{D}}_{\text{Direct}} or DCoT{\mathcal{D}}_{\text{CoT}} (discussed below) of problems-solution pairs with correctness tokens (optionally with CoT rationales), GenRM minimizes the loss:

where λ>0\lambda>0 is a hyperparameter that controls the data mixture ratio between solution verification (Dverify{\mathcal{D}}_{\text{verify}}) and generating correct solutions (Dcorrect{\mathcal{D}}_{\text{correct}}). This unified training can improve verifier and generation performance via positive transfer between these two related tasks: how to generate a correct solution, and whether a solution is correct. By default, we train GenRM verifiers using the unified loss in (5).

3 Chain-of-Thought Verifiers (GenRM-CoT)

Since verification often involves complex reasoning, generative verifiers can naturally benefit from CoT reasoning (Wei et al., 2022). Specifically, we can generate intermediate reasoning steps or critique (CoT) before making a decision about the solution correctness, which may identify subtle reasoning errors missed by direct verifiers (Figure 3, bottom). To train CoT verifiers, we can minimize the SFT loss LGenRM{\mathcal{L}}_{{\text{GenRM}}} on the dataset DCoT{\mathcal{D}}_{\text{CoT}} containing problem-solution pairs as inputs, and corresponding verification rationales vCoT{\mathbf{v}}_{\textbf{CoT}} appended with a final question I{\mathbf{I}} and ‘Yes’ or ‘No’ token as targets:

Notably, these rationales can either be human-generated or LLM-generated, both of which we explore in this work. During inference, we first generate a CoT rationale vCoT{\mathbf{v}}_{\textbf{CoT}} from GenRM-CoT and then use the probability of ‘Yes’ for assigning the correctness score:

Compared to (4) that directly uses the instruction I{\mathbf{I}} to produce a score, we can see that the above CoT reward additionally conditions on ICoT{\mathbf{I}}_{\textbf{CoT}} and self-generated vCoT{\mathbf{v}}_{\textbf{CoT}} before getting a score via instruction I{\mathbf{I}}.

Inference-time Compute for CoT verifier When sampling verification CoTs, the generative verifier may use different reasoning paths and yield different correctness probabilities for the same problem-solution pair. As such, we would like to marginalize out the intermediate reasoning paths to select the most consistent correctness answer (Wang et al., 2022). To do so, we can use majority voting where we first generate KK verification CoT rationales, and average the CoT-verifier score for these rationales:

Since individual verification rationales from CoT verifiers can have reasoning errors, majority voting can mitigate the impact of such errors by averaging correctness scores across multiple rationales. Importantly, this means that GenRM-CoT can leverage additional inference-time compute to improve its accuracy, which discriminative verifiers cannot do. Unless otherwise specified, we report GenRM-CoT performance based on majority voting with 32 votes, that is, K=32K=32 in (7).

Synthetic Verification Rationales for CoT Verifier. Verifying LLM solutions with human-generated rationales can become increasingly expensive and challenging as LLMs surpass human reasoning abilities. To address this challenge, we explore using synthetically-generated rationales on GSM8K. One naive approach is to simply use the ‘Let’s verify step by step’ prompt given a problem-solution pair, and filter the generated rationales based on whether they correctly verify the correctness of a solution. However, such rationales are often of poor quality due to 50% accuracy from random guessing.

Instead, we use reference-guided grading (Zheng et al., 2024) to improve the quality of synthetic rationales. Specifically, we provide a reference solution in addition to the problem and solution to verify (see Table A.2), making it easier for an LLM to point out any reasoning error in the provided solution. Here, a reference solution is defined as any model-generated solution that manages to arrive at the correct final answer. Note that reference-guided grading can only be used during training, as we do not have reference solutions for test problems.

Experiments

The goal of our experiments is to demonstrate the efficacy of next-token prediction compared to other approaches for training verifiers. To this end, we compare GenRM and standard verifiers on a number of reasoning problems with the goal of answering the following research questions:

How does GenRM compare to standard discriminative verifiers in reasoning domains?

Does unified training of GenRM improve generation and verification performance?

Can GenRM effectively utilize CoT reasoning and test-time compute to improve its performance?

How does GenRM scale with data and model size?

Tasks and Data Generation. We focus on the following reasoning tasks (more details in Appendix A):

Last Letter Concatenation (Wei et al., 2022): Given a list of words, the task is to concatenate the last letters of each word (for instance, "Noah Paul Elisha Rebecca" →\rightarrow "hlaa"). We train verifiers on lists of lengths 2−42-4, and evaluate the verifier on the out-of-distribution (OOD) setting of length 6.

Word Sorting (Suzgun et al., 2022) Given a list of words, sort them in alphabetical order. For a sorting task, verification is clearly easier than generating the correct solution. Similar to the last letter concatenation task, we train verifiers on up to 4-words, and evaluate length-generalization performance on 5 word examples.

GSM8K (Cobbe et al., 2021) is a widely-used dataset to evaluate grade-school math reasoning capabilities of LLMs. For training, we use at most 16 correct and 16 incorrect solutions per problem. We evaluate the verifier performance on 16 solutions per problem in the test set.

Baselines. We compare GenRM to the following standard approaches for verification:

Discriminative (standard) RM (ORM) (Cobbe et al., 2021): The prevalent approach for training verifiers for test-time re-ranking on reasoning tasks, as discussed in §2, serves as our main baseline.

Self-consistency (Wang et al., 2022): A simple approach to use test-time compute without verifiers: sample multiple solutions from the LLM generator and pick the most common answer.

LLM-as-a-Judge (Zheng et al., 2024): This approach uses an off-the-shelf pretrained LLM for verification. To do so, we use a CoT prompt to produce 32 verification rationales that is used for correctness prediction and pick the majority vote correctness answer.

Evaluation protocol. Following prior work (Cobbe et al., 2021; Lightman et al., 2023), we primarily use Best-of-N performance (in terms of the percentage of problems solved) of a fixed LLM (generator) (§2) with learned verifiers, and report average performance on the test set. Best-of-N evaluates what fraction of solutions chosen by the verifier are correct. For some results, we also report RM accuracy on the test set or a held-out validation set, which measures whether the verifier accurately classifies incorrect and correct solutions. While these two metrics are often correlated, RM accuracy only evaluates the verifier’s point-wise accuracy, while Best-of-N evaluates its group-wise ranking performance.

Models. For training discriminative RMs and GenRM, we use open-weights Gemma models (Team et al., 2024), specifically Gemma-2B for algorithmic tasks, and Gemma 2B, 7B, and 9B for GSM8K. For solution generation as well as LLM-as-a-Judge, we use Gemma 2B for algorithmic tasks and Gemini 1.0 Pro (Team et al., 2023) for GSM8K.

Hyperparameters. By default, all GenRM experiments use unified training for verification with solution generation (5), with λ=1/3\lambda=1/3 for algorithmic tasks and λ=1/4\lambda=1/4 for GSM8K. We use the label ‘Verification Only’ to indicate GenRMor GenRM-CoT verifiers trained using only verification data (λ=0\lambda=0). See Appendix Appendix B for more details.

Verification CoT Rationale Generation. For training rationales, we generated ground-truth rationales for word sorting and last letter concatenation algorithmically, as shown in Table A.1. On GSM8K, we generated rationales using Gemini 1.0 Pro with reference-guided grading (Zheng et al., 2024), with the prompt outlined in Table A.2.

GenRM, which directly predicts Yes/No token for verification, can match or outperform the discriminative RM and other approaches on all the three tasks, as shown in Figure 6. This shows that the next-token-prediction loss allows GenRM to tap into the capabilities of pretrained Gemma models more effectively.

CoT Reasoning Improves Verification. GenRM-CoT, which combines chain-of-thought with majority voting, further improves the performance over GenRM.

In particular, on the two algorithmic tasks with oracle verification CoTs, GenRM-CoT closely matches the oracle verifier performance (Pass@N). On GSM8K, GenRM-CoT consistently outperforms all other methods, even though the training CoT rationales (generated with Gemini 1.0 Pro) may contain errors. Qualitatively, GenRM-CoT is able to detect subtle reasoning errors that are missed by discriminative verifiers (see Figure 2, 5,4).

2 Unifying Generation and Verification

Unifying solution generation with verification, as done by GenRM using the next-token-prediction objective, consistently improves verification performance across all tasks, as illustrated in Figure 7. This improvement is observed for both direct and CoT-based generative verifiers, suggesting that teaching the verifier to imitate correct solutions generally helps.

Notably, incorporating CoT verification data into the generator’s training mix leads to better solution generation performance for the GenRM-CoT verifier itself, as evidenced in Figure 8 by the improved Best-of-N scores with the oracle verifier (Pass@N). This suggests that teaching a generator to perform verification based on next-token prediction can deepen its understanding of the generation process itself. Overall, the above results indicate the unifying solution generation and verification is mutually beneficial.

3 Scaling Data, Model Size, and Inference-time Compute

Scaling Test-Time Compute with GenRM-CoT can be done by sampling multiple CoTs and conduct majority voting, as described in (7). As shown in Figure 9, GenRM-CoT verifier’s performance scales gracefully with greater number of votes at test time, under all three Gemma scales (2B, 7B, 9B), outperforming greedy decoding performance within 4 votes. Across model scales, the finetuned GenRM-CoT verifier outperforms LLM-as-a-Judge, which also utilizes the same CoT approach and number of majority votes, but prompts a more capable Gemini 1.0 Pro model than Gemma models.

Scaling model size. In Figure 10, we show that the performance of generative verifiers scales up positively with an increase in model capacity. The model scaling experiments use the fixed dataset (32 solutions per problem, and for CoT GenRM, 4 rationales per solution). The results show that bigger models are able to learn more from the same data, which matches what we expect from scaling model parameter counts under the next-token prediction loss.

Data scaling for CoT verifiers. GenRM-CoT allows for an additional axis of data scaling, which is absent in standard verifiers: scaling the number of rationales per solution. On GSM8K, we find that using multiple rationales per solution has a substantial effect on the performance of generative verifiers. We suspect that this is because model-generated, synthetic rationales are noisy in this case, such that training on multiple rationales per solution and the associated “ensembling” effect prevents the training procedure from overfitting to this noise and spurious correlations. See Figure 11 for detailed comparisons. Both RM Accuracy and Best-of-N Accuracy scales positively on both data axis, with the number of rationales per solution having a bigger effect.

Data scaling for direct GenRM. GenRM trained on verification only data still outperforms standard verifiers as we increase the number of solutions per problem on GSM8K (Figure 13, left), demonstrating the effectiveness of casting verification as a next-token prediction problem. Across all data scales, unified training with solution generation data further boosts the performance of GenRM verifiers, as already discussed in §4.2. Moreover, the optimal loss coefficient for solution generation data in GenRM follows an inverted U-shape: adding too little or too much negatively impacts verification, while intermediate values yield the best results (Figure 13, right).

4 Impact of Synthetic Rationale Quality

Our results on GSM8K results indicate that GenRM-CoT verifier can outperform discriminative and direct GenRM verifiers even without human-written rationales, highlighting the potential of LLM-generated rationales. However, the quality of these synthetic rationales does matter, as shown in Figure 13. Using reference-guided grading during rationale generation significantly improves performance (91.7% with guidance vs. 87.8% without for Gemma-7B verifiers), indicating that LLMs are better at identifying reasoning errors when they have a reference solution for comparison. Importantly, achieving our result does not require a more capable model for generating verification rationales: we use the same model (Gemini 1.0 Pro) to generate both solutions to verify and synthetic rationales in the training data.

Related Work

Reward models (RMs) and verifiers. Conventionally, RMs and verifiers are trained as discriminative models via binary classification: given a prompt and a corresponding solution or a pair of solutions), the model is either trained to predict the correctness of the solution (Cobbe et al., 2021; Saunders et al., 2022; Lightman et al., 2023; Wang et al., 2023; Uesato et al., 2022; Luo et al., 2024; Yu et al., 2024) or a preference between the two solutions (Stiennon et al., 2020; Nakano et al., 2021). Concretely, the RM or verifier directly produces a numerical continuous-valued score, which is then plugged into a classification objective (3). As such, discriminative verifiers do not utilize the text generation capabilities of LLMs. In contrast to discriminative RMs, GenRM does not train RMs that output a numerical score from a special logit, but rather represent the correctness decision using the log probability of ‘Yes’ and ‘No’ tokens under special instructions. Posing verification as generating “yet another token” allows it to tap better into the generation capabilities of LLMs, by making it straightforward to employ CoT reasoning and additional inference-time compute for better verification.

LLM-as-a-Judge for verification. Another line of work that poses verification as next-token prediction simply prompts off-the-shelf LLMs to act as a verifier when provided with a rubric and a template for grading (Zheng et al., 2024; Bai et al., 2022; Kim et al., 2023; Ling et al., 2024), but without any specific training for the same. Perhaps unsurprisingly, we find in our experiments that using substantially more powerful LLMs (Gemini 1.0 Pro) as a judge is substantially worse than our trained GenRM (using weaker Gemma models), highlighting the necessity of training generative verifiers. More generally, even the strongest proprietary LLMs to date, such as GPT-4 (Achiam et al., 2023) and Gemini (Team et al., 2024b), fall behind trained RMs on popular leaderboards, such as RewardBench (Lambert et al., 2024), and this gap is much larger for reasoning problems. While Agarwal et al. (2024) utilized many-shot prompting on an off-the-shelf LLM to obtain a “Yes/No” verifier, our work focuses on training generative verifiers and also leverages CoT rationales and test-time compute.

Using CoTs for reward models. Prior works have also considered using critiques or CoT to extract preference and verification signals (Yuan et al., 2024; Wu et al., 2024; Wang et al., 2024; Ye et al., 2024a); in contrast to these prior works, GenRM utilizes model-generated CoT directly for training the verifier. Upon inference, a GenRM produces its own CoTs, which it then uses to make decisions on correctness, unlike other work that simply uses CoTs from a separate (and highly-capable) model (Ye et al., 2024b). Compared to (McAleese et al., 2024) which assumes access to high-quality data from humans to train discriminative RMs for generating code critiques, we show that GenRM can be trained from purely synthetic, model-generated critiques.

Concurrent work (Ankner et al., 2024) trains a model to produce response critiques, which are then passed as input into a reward modelling head, separate from the base LLM. Unlike GenRM which uses next-token prediction, their RM head is trained discriminatively akin to standard RMs. While this approach allows them to leverage CoTs and inference-time computation, it does not allow them to unify solution generation and verification together as a result of a discriminative RM head, that GenRM seamlessly enables (Section 4.2). Moreover, our results show that even using GenRM to produce a single verification token outperforms standard RMs even without CoT or inference-time compute, which can potentially improve the results in this concurrent work (Figure 13, Left).

Unified generation and verification. One of the hallmark properties of GenRM is that the very same generative verifier can be co-trained with a generation objective (5): when given a problem, the model is trained to produce a solution, whereas when given a problem and a candidate solution, it is trained to verify this candidate. This is related to DPO (Rafailov et al., 2024) and its application to learning verifiers in reasoning (Hosseini et al., 2024), which aims to unify generation (policy) and verification (reward models) by representing the reward implicitly using the logits of a policy and training the policy with a reward-modelling loss. For reasoning, this type of model tying has been shown to exhibit erroneous extrapolation and a degradation in learned representations, which prior work has attempted to address with additional techniques (Pang et al., 2024; Setlur et al., 2024; Pal et al., 2024; Yang et al., 2024). Of these, while the approach of Yang et al. (2024) trains an reward model with an auxiliary generative SFT loss, note that this loss is applied on a separate head for regularization purposes and is discarded after training; unlike GenRM no text is produced when querying the RM.

Unlike DPO, our unification approach is distinct: GenRM treats verification as next-token prediction, allowing the same model to be trained on both verification and policy tasks using their respective prompts. This eliminates the need for separate reward modeling or loss modifications, offering greater flexibility for practitioners to combine various reward and generation objectives. For example, while Hosseini et al. (2024) struggled to obtain a DPO verifier that can generate correct solutions, the GenRM-CoT verifier has better generation performance compared to SFT only on correct solutions (Figure 7).

Conclusion & Future Work

In this paper, we have introduced Generative Verifiers (GenRM), which recast verification as next-token prediction in LLM reasoning domains. GenRM is a more performant alternative to discriminative reward models, and unlocks the use of powerful tools, such as chain-of-thought reasoning and majority voting for better verification. GenRM also unifies generation and verification into a single LLM, and demonstrates that such a unification benefits both generation and verification. Moreover, we show that synthetic model-generated rationales, which are noisy and sub-optimal in many aspects, are sufficient to teach GenRM how to use verification CoT to pick out tricky reasoning errors in grade school math problems. Future work includes extending the generative verification framework to a broader range of tasks (such as coding, alignment, and answering open-ended questions), and studying how generative verifiers can be integrated into existing LLM self-improvement algorithms (Singh et al., 2023; Gulcehre et al., 2023).

As a generative LLM for verification, GenRM offers a solid foundation for future work. It’s readily compatible with all the existing tools designed to improve LLMs, such as retrieval-augmented generation (Borgeaud et al., 2022), many-shot learning (Agarwal et al., 2024), multi-staged prompting (Yao et al., 2024), tool use (Schick et al., 2024), and code generation and execution (Chen et al., 2021). Future work can explore how those techniques can be in service of verifying model-generated outputs under the generative verifier framework.

Acknowledgements

This work was done during LZ, AH, and HB’s internship at Google. We thank Hugo Larochelle, Minqi Jiang, Aleksandra Faust, Ankit Anand, Guillaume Desjardins, Doina Precup, and Charlie Snell for feedback on an earlier version of this paper and informative discussions. We thank Chirag Nagpal and Katrin Tomanek for support in setting up infrastructure for experimentation that was crucial for running our experiments.

References

Appendices

Last Letter Concatenation: To generate the training data, for each length {2,3,4}\{2,3,4\}, we generate 350350 problem queries by randomly sampling from the set of words in original training set; for each problem query, we generate 128 attempts from Gemma-2B (Team et al., 2024a) model. This gives us a total of about 50K training data points after de-duplication. We train verifiers on examples of lengths {2,3,4}\{2,3,4\} (here the length refers to how many words are in the input list), and evaluate the verifier performance on length 6. We use the format in Table A.1 to algorithmically generate ground-truth verification CoT for training.

Word Sorting: We train verifiers on a dataset comprised of {2,3,4}\{2,3,4\} words in each example, and evaluate the performance on length 55. For each length, we generate 4096 lists of words as the problem queries; for each problem, we generate 64 attempts from Gemma-2B. After de-duplication and filtering out invalid responses, we have a total of about 100K training data points. We also algorithmically generate ground-truth verification CoT for training (see Table A.1).

Grade School Math: We follow the original train/test split and use 1.3K problems for test, 128 problems for validation, and about 7.2K problems for training. We generate 50 solutions per problem, and randomly sample at max 16 correct solutions and 16 incorrect solutions per problem as the training set.

Appendix B Hyper-parameters for Verifier Training

We finetune Gemma-based language models. After doing a sweep of learning rates (LR), we find that an LR of [2e−6,1e−6,5e−7][2e-6,1e-6,5e-7] works well for our tasks considered (with LR=2e−62e-6 generally being the best). We use a weight decay of 1e−21e-2, and do not apply any dropout. We use the Adam optimizer (Kingma, 2014) with decoupled weight decay (Loshchilov and Hutter, 2017) and a gradient norm clipping of 1.01.0. We use a linear warmup of 10001000 gradient steps, and a cosine decay schedule that decays to 10%10\% of the peak learning rate towards the end of training. We finetune for 200K steps with a batch size of 128, and use seqio (Roberts et al., 2022) library to create data mixtures. We pick the best checkpoint based on validation accuracy of verification on held out problems and solutions. We always use data balancing between 50% correct solutions and 50% incorrect solutions in training.

We finetune Gemma-based discriminative RMs by using a special token’s logit for classification. We chose the best performing ORM on our validation sets by launching a large sweep over learning rates [1e−7,5e−7,1e−6,2e−6,3e−6,5e−6][1e-7,5e-7,1e-6,2e-6,3e-6,5e-6], weight decay [1e−3,1e−2,1e−1][1e-3,1e-2,1e-1] and dropouts [1e−3,1e−2,4e−2][1e-3,1e-2,4e-2]. We also schedule the learning rate with a linear ramp up and a cosine decay. We then pick the best model according to validation accuracy of verification on held out problems and solutions. We always use data balancing between 50% correct solutions and 50% incorrect solutions in training.

Appendix C Examples of verification rationales generated by GenRM-CoT on GSM8K