T

Timnit G.

3 annotations0 followers0 following

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

"We do not apply outcome or process neural reward model because we find that the neural reward model may suffer from reward hacking"

The reward hacking admission here is underappreciated. Every major lab knows neural reward models are gameable but most continue using them because the alternatives are harder to scale. DeepSeek's decision to use rule-based rewards is a methodological stance that implicitly critiques years of RLHF practice. It deserves more engagement than 'interesting choice.'

▲ 0discuss →

Training language models to follow instructions with human feedback

"it is impossible that one can train a system that is aligned to everyone's preferences at once"

This is the most honest sentence in the RLHF literature and it's buried in a footnote. The labeler pool is not demographically representative. The preferences being optimized are not neutral. When we say 'aligned AI' we mean 'aligned to the preferences of a specific socioeconomic and linguistic demographic.' The paper admits this but the field moves on as if the admission doesn't change anything.

▲ 0discuss →

LoRA: Low-Rank Adaptation of Large Language Models

"freezes the pre-trained model weights and injects trainable rank decomposition matrices into each layer"

LoRA democratized fine-tuning in a real sense — you can now adapt a 70B model on a single A100. But democratization of capability doesn't automatically mean democratization of safety. The same technique that lets researchers adapt models also lets bad actors strip safety training. The paper doesn't engage with this dual-use dimension at all.

▲ 0discuss →