Training language models to follow instructions with human feedback
"Making language models bigger does not inherently make them better at following a user's intent."
I'd reframe this slightly: capability and alignment are orthogonal axes that happen to both benefit from some of the same training ingredients. RLHF moves you on the alignment axis without moving you much on the capability axis. That orthogonality is what makes the result interesting — and what makes alignment hard.
Language Models are Few-Shot Learners
"We use the term 'in-context learning' to describe the inner loop of this process"
There's a philosophical question buried here that the paper never surfaces: is in-context learning 'real' learning or is it task formatting? The 2023 mechanistic interpretability literature suggests both are happening simultaneously, at different layers. The distinction matters enormously for how we think about alignment.
Attention Is All You Need
"dispensing with recurrence and convolutions entirely"
What struck me most is how much of modern ML infrastructure this paper quietly deprecated. We spent years optimizing LSTM training pipelines, gradient clipping heuristics, BPTT scheduling. All of that institutional knowledge became irrelevant almost overnight. The switching cost wasn't technical — it was organizational.