Outcome-based Reinforcement Learning to Predict the Future
Benjamin Turtel, Danny Franklin, Kris Skotheim, Luke Hewitt, Philipp Schoenegger
Introduction
Reinforcement learning with verifiable rewards (RLVR) has recently boosted large-language-model (LLM) performance on benchmarks such as GSM8K, AIME, and MATH(Zhao et al., 2025a; Shao et al., 2024; Guo et al., 2025; Lambert et al., 2024; Xie et al., 2025). RLVR fine-tunes a base model with on-policy reinforcement learning (RL) against objective pass/fail signals, like unit tests or exact numeric solutions(Liu et al., 2023; Su et al., 2025). Complementary self-learning work shows that mixing verifiable non-math corpora into RL pipelines can extend these gains beyond purely symbolic tasks(Akter et al., 2025).
However, it remains unclear how well such improvements transfer to messy, real-world reasoning problems whose outcomes are noisy, delayed, and probabilistic rather than deterministic and instantly verifiable(Murphy, 2025). Forecasting exemplifies this challenge: it demands causal inference, trend extrapolation, and well-calibrated probabilities while supplying only sparse, lagged supervision(Weng, 2024; Qu et al., 2025). Under these conditions standard GRPO-style updates can drive policies toward extreme overconfidence, gibberish output, or outright training collapse. Adapting RLVR to this setting promises to extend model reasoning ability to an especially demanding domain, potentially unlocking a host of real-world applications.
We design and empirically validate a stable RL pipeline for probabilistic forecasting. On the algorithmic side, we (i) remove per-question standard-deviation scaling from Group Relative Policy Optimisation (GRPO), (ii) switch to baseline-subtracted advantages in ReMax(Li et al., 2023), and (iii) add lightweight guard-rails, token-length limits, a gibberish filter, and an early-stop criterion, to keep gradients proportional to Brier loss and prevent collapse over more than 100,000 sequential events. On the evaluation side, we measure accuracy with the soft-Brier score and calibration with expected calibration error (ECE) on a strictly time-ordered 3,300-question Polymarket hold-out, and we quantify economic value by converting each forecast into a set of hypothetical trades and comparing realised profits with those of the frontier reasoning model o1 as a benchmark.
We find that on the 3 300-question hold-out set, a seven-run ReMax ensemble attains a Brier of and an ECE of , which is not statistically different from o1 in accuracy (, ) while halving its mis-calibration (, ). With a trading strategy that only places a hypothetical bet when the model’s predicted edge exceeds its own expected-calibration error, ReMax realises \127\ for o1, being restricted to a single \ is statistically significant (95 % CI +\2.0+\, ). Thus our ReMax ensemble is on par with frontier model accuracy, reaches near-market-level calibration, and leverages that calibration edge in a hypothetical trading strategy to outperform o1.
Method
We construct a time-ordered corpus by merging two sources of questions. First, we collect roughly 10,000 resolved yes/no contracts from Polymarket, including creation date, close date, resolution timestamp and final outcome. Second, we generate 100,000 additional forecasting prompts with Lightning Rod Labs’ proprietary Foresight Learning framework, generating high-quality training data without humans in the loop or labeled data. The final dataset covers multiple topical domains ranging from macroeconomics, weather, and culture to technology and politics. Each question is binary (resolves as yes or no) and has a probability attached to it (market price or model forecast). All training examples resolve before the earliest test question, fully eliminating look-ahead bias. For every contract we draw a single “prediction timestamp” uniformly at random between its on-chain open and scheduled close; all external context is truncated at 00:00 UTC on that calendar day. News headlines that we provide the models with at prediction time are retrieved through the Exa.ai API, which offers day-level granularity, so the model never sees information dated on or after its own forecast.
We first train a set of models with the initial 10,000 question set. In a second set of experiments, we use the same initial 10,000 questions with the addition 100,000 synthetic questions mixed in and time-ordered. We reserve a held-out test set of 3,300 questions for all models.
2 Model
All experiments are conducted with DeepSeek-R1-Distill-Qwen-14B, a 14-billion-parameter model initialised from the open-weight Qwen 2.5-14B base checkpoint (Yang et al., 2024) and further instruction-tuned on the 800k-sample DeepSeek-R1 distilled reasoning corpus(Zhao et al., 2025b). Our fine-tuning regimes take this as the base model and our analyses compare to this base model as well as other benchmarks and frontier models.
3 Reward Design
We treat the forecasting task as a sequential decision problem in which the model receives a textual prompt about a future event, outputs a single probability that the event will occur, and subsequently observes the binary outcome . The episode reward is the (negative) Brier score(Brier, 1950)
a strictly proper scoring rule that incentivises calibrated probabilistic forecasts.
Model outputs occasionally fail to parse as a valid probability, for example when the generation omits a numeric answer or when the model fails to answer altogether. During training, we use the following scheme:
Strict Brier: assign the maximum Brier loss of (equivalently a training reward ) whenever the regex fails to extract a probability in $$.
For evaluation, we use the following metric that does not penalise misformatting and missing predictions as much:
Soft Brier: assign a constant Brier loss of (training reward ) to any malformed or missing forecast, preserving gradient signal while avoiding training collapse.
We evaluate three on-policy RL algorithms, GRPO, Modified-GRPO and ReMax, and include Direct Preference Optimisation (DPO) as an off-policy baseline.
Second, we modify the standard GRPO algorithm by removing the standard-deviation division and set . This preserves the raw magnitude of especially large forecast errors, potentially improving the model’s ability to correct extreme miscalibrations. Yet, omitting the normalisation can make optimisation more sensitive to outliers, requiring additional guard-rails to mitigate instability.This vulnerability is partly mitigated in our setting because Brier scores are intrinsically bounded in $$, which caps the variance of individual rewards and makes the absence of per-question normalisation far less destabilising than it would be for tasks with unbounded numeric losses. This modification parallels the Dr. GRPO method proposed by Liu et al.(Liu et al., 2025).
where balances a KL penalty term. Subtracting a baseline in lieu of variance normalisation often better preserves large reward signals.
Lastly, we also test Direct Preference Optimization(Rafailov et al., 2023), a direct preference-based approach. This method has shown strong performance on QA tasks as well as forecasting tasks (Turtel et al., 2025) and functions as a baseline for some of our analyses.
In our forecasting context, normalising each question’s rewards by their standard deviation (as in standard GRPO) can excessively dampen large errors and encourage overconfidence: per-question normalisation flattens the reward distribution and erases the natural asymmetry whereby modest gains accrue from correct but overconfident predictions, while rare mis-predictions incur disproportionately large penalties, the very signal the model needs to learn proper calibration. By removing per-question normalisation (Modified GRPO) or using ReMax’s baseline-based approach, we better preserves the impact of large deviations, improving calibration when probabilistic forecasts deviate significantly from eventual outcomes.
Across algorithms we keep the optimisation scaffold identical: AdamW (, , , no weight decay), bfloat16 precision, global grad-norm clip , and an entropy bonus coefficient of . We modify only the levers each method cares about. GRPO uses an actor learning rate of , an initial KL penalty of , PPO ratio-clip , and roll-outs per prompt; Modified-GRPO is identical except that it drops the division in the advantage to isolate the effect of normalisation. ReMax doubles the actor learning rate to , keeps the KL schedule unchanged, and trains its learned value baseline with under an MSE loss scaled by . DPO is run for epochs at with a constant learning rate and a batch size of sequences. All runs employ automatic mixed precision and gradient accumulation to emulate two sequences per GPU.
All experiments ran on a single 8-GPU node. The GRPO (10k), ReMax, Modified-GRPO, DPO, and the large-scale GRPO-100k run used eight NVIDIA H100 GPUs.
4 Training Protocol and Stability Measures
All on-policy algorithms (GRPO, Modified-GRPO, ReMax) are trained strictly online: each question is encountered only once in chronological order, and its outcome is revealed immediately after the event date. We do not perform multiple epochs, as re-exposing the model to past questions after outcomes are known leads to severe over-fitting (the model essentially “learns the future” on subsequent epochs).
While our primary focus is on comparing the four RL algorithms, we also benchmark forecasts against two additional references:
OpenAI’s o1, prompted with the same question text, to assess performance against a frontier reasoning model.
Market prices (Polymarket), the market’s implicit probability at the time each question was asked. For every test question, the Polymarket price is sampled at precisely the same timestamp used for the prompt cut-off, yielding a strictly contemporaneous benchmark.
These baselines indicate how our RL-trained models perform relative to both a state-of-the-art commercial LLM and a real-world prediction market.
5 Failure Modes
We observed a set of different failure modes. For example, in the model trained with GRPO without any guardrails, we find that the model sometimes assigns a probability of 0 to outcomes that are not impossible but rather unlikely. Asigning a probablity of 0 is not something a capable forecaster would do and should be avoided by forecasting models. For example, in response to the question “Will Trump say “Ukraine” 20+ times during Macron presser today?”, the model outputs
[…] In conclusion, while Ukraine will definitely be a topic, the number of mentions reaching 20 is astronomically low. Even if every statement and answer included "Ukraine," it’s improbable to reach that number. Therefore, my estimate is that the probability is extremely low, almost zero. *0.000*.
Similarly, the model shows this overconfidence also on the other end of the probability spectrum. While acknowledging that it may have misunderstood something, it assigns probability of 1 to questions that the model itself is somewhat unsure about: “Therefore, unless I’m missing something, the probability is 1.” Overall, this model has 39.3% of predictions in either the 0–10% or the 90–100% bucket, with many landing exactly on 0 and 1, which shows a prevalence for very high or low probabilities.
Using our Modified GRPO approach, again without any guardrails, we do find that the model is more likely to avoid the extreme overconfidence of the standard GRPO approach, with only a total of 7.9% of predictions landing in either of the extreme buckets. However, we do observe a different failure mode. Some of the answers show language switching between Chinese and English within an output, “Alright, so I need to预测 Patrick Mahomes 在超级碗 LXIX 中是否会 rush 30+码,结果以超过30码为“Over”,否则为“Under”。首先,我会回顾提供的新闻和背景信息,找出相关数据和趋势”, while others show partial segments of ungrammatical phrases and gibberish, such as “Next, I look at the …think that would happen in this interview,” “I s…,” “Lamar Jac…that,” and “They talk about the theme, ticket distribution, and guests, but t…that they’re involved.”
When training the model with our set of guardrails in place on the 100k data set, both types of failure modes are less pronounced. For example, 13.1% of predictions fell into the 0–10% and 90–100% buckets, with many forecasts that were close to 0% now having values that are greater than 0. Similarly, gibberish responses are less likely, though some ungrammatical phrases and elisions remain, such as in “Looking a…ould impact the price.”
Results
We measure predictive accuracy with the soft-Brier score, defined as the squared error averaged over all 3,300 questions but assigning a soft penalty of whenever a model fails to produce a parseable probability. A score of 0.25 is equivalent to guessing 50% on a question, and thus functions as a suitable stand-in for failed responses. For calibration, we use the expected calibration error (ECE) computed in ten equal-mass probability bins. For both accuracy and calibration, lower scores indicate higher accuracy and better calibration respectively. All statistics are paired across the identical question set (); confidence intervals (CI) are two-sided Wald intervals for Brier and bootstrap intervals for ECE unless specified otherwise; every -value reported below is two-sided.
Among the 10k-trained models, ReMax (no guard-rails) is the most accurate with a soft-Brier score of 0.197 [0.190, 0.204]. DPO follows at 0.205 [0.197, 0.213] (), then Modified GRPO at 0.206 [0.199, 0.213] (), vanilla GRPO at 0.210 [0.200, 0.220] () and the base model at 0.214 [0.207, 0.221] (). Calibration shows the same ordering: ReMax posts an ECE of 0.0507 [0.0381, 0.0633]; the nearest competitor, Modified GRPO, is higher by 0.046 ECE points [0.0325, 0.0599], . DPO, vanilla GRPO, and the base model are still farther off, each differing from ReMax by at least 0.035 ECE points with .
Figure 1 illustrates that ReMax, even at the 10k scale, closes roughly one-third of the accuracy gap between the instruction-tuned base model (Brier 0.214) and the real-money Polymarket benchmark (0.162) and delivers the most reliable probability forecasts among all training algorithms evaluated.
2 Main Model Accuracy Comparison.
Table 1 reports mean soft-Brier score and expected-calibration error (ECE) for the 7-run ReMax ensemble and Modified-GRPO (both trained on the 100k corpus) alongside the instruction-tuned base model (DeepSeek-R1 14B), the frontier baseline OpenAI o1, and contemporaneous Polymarket prices.
Relative to the base model, ReMax lowers Brier by points () and reduces ECE by roughly a third (, ). Its Brier is statistically indistinguishable from o1 (, ), while its calibration is markedly better ( ECE, ) and essentially matches Polymarket ( ECE, ). Modified-GRPO also improves substantially over the base model ( Brier, ; ECE, ), but remains slightly worse than ReMax (Brier gap , ) and shows no reliable difference to o1 (Brier gap , ; ECE gap , ). Both learning schemes still trail the real-money market in Brier (ReMax , Mod-GRPO ; ), yet ReMax shows very strong calibration.
3 Hypothetical Trading Evaluation
To evaluate our forecasting algorithms in a more direct test, we convert every probability into a hypothetical one-share trade against the contemporaneous Polymarket price. For each resolved contract with non-zero volume we compare the model’s probability with the last quoted market price . If the strategy buys one m+0.01p
Figure 2 visualises how each model converts forecast skill into cash. The left panel shows cumulative realised profit as we step through the 3132 Polymarket questions in descending order of expected edge; the right panel condenses the final take under the three bet-selection rules.
In the high-edge analysis (Edge ECE) ReMax secures the largest absolute return, netting \127\, OpenAI o1 with \92\. Uncertainty was quantified with a question-level paired bootstrap (9999 replicates): each replicate resampled the set of Polymarket questions with replacement, kept the four models’ profits for every question together (thereby preserving their correlation), summed those profits to obtain a total for each model, and then formed percentile 95 % confidence intervals; two-sided -values were obtained by centring the bootstrap differences at zero. On this analysis, ReMax out-earns o1 by +\35.6+\ to +\68.6p=0.037+\ (CI +\23.3+\, ), while its advantage over Mod-GRPO, +\16.9-\ to +\42.3p=0.19+\ (CI +\2.7+\, ), whereas the Mod-GRPO vs. o1 gap (+\18.7-\ to +\56.4p=0.33+\, CI -\8.8+\, ) remain inconclusive. Overall, ReMax clearly outperforms both o1 and DeepSeek, Mod-GRPO also bests DeepSeek, and the remaining contrasts are indistinguishable within sampling error.
When trading only on high signals (Edge ECE), the models keep roughly 60-80 % of the wagers that show any positive edge (ReMax 2 193 / 2 931, Mod–GRPO 2 353 / 2 985, o1 1 718 / 2 935, DeepSeek 1 820 / 3 013) yet harvest almost the entire per-trade profit. In this band ReMax averages 5.8 ¢ per trade (95 % CI ¢) for a total of \127[3.1,\,6.3]\; o1 5.4 ¢ and \92[2.2,\,5.7]\. Pair-wise Welch tests on these per-trade means find no significant differences among ReMax, o1 and Mod–GRPO (largest gap = 1.1 ¢, ), indicating that their high-edge per-trade performance is statistically indistinguishable; ReMax’s 1.8 ¢ advantage over DeepSeek is also not significant ().
Applying a simple calibration filter, placing a bet only when the model’s predicted edge exceeds its own ECE, retains about 80% of the trading opportunities yet harvests at least 95 % of the attainable profit. Within this high-edge regime, ReMax stands out as the top earner, beating both o1 and the DeepSeek base model in total hypothetical profit. Mod–GRPO and o1 deliver broadly comparable returns, and all three fine-tuned models comfortably exceed the untuned DeepSeek baseline. In short, calibration gating captures most of the dollars on the table, and the ReMax ensemble, due to its strong calibration, is the best of our models to deploy under that rule.
Figure 3 relates Polymarket’s own confidence to the ex-post betting edge of the ReMax Ensemble-7. The horizontal axis spans (complete uncertainty) to (full certainty). The vertical axis shows the model’s mean excess win probability in percentage points (pp), with a dashed zero line marking break-even. We fit a penalised cubic-regression spline via mgcv::gam (REML, effective ) and plot it as a solid curve, while the shaded area denotes the pointwise confidence band. Tick labels above zero carry explicit plus signs, and axes are clipped to the empirically populated rectangle .
The advantage of ReMax is concentrated in low-confidence markets and vanishes as certainty rises. This could be due to cases where the market has more up-to-date or more accurate information than our model, which has an information lag to ensure never “cheating” during market simulation and may miss especially relevant information. In the confidence band the mean edge is (t-test, ); for moderate confidence it declines to () but remains positive. Above (80%) market confidence the residual is essentially zero (, ).