DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
"Reinforcement Learning"
The RL approach here is notably different from RLHF — it uses pure outcome-based reward without a learned reward model.