← paper
Attention Is All You Need

Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, Illia Polosukhin

“dispensing with recurrence and convolutions entirely”

What struck me most is how much of modern ML infrastructure this paper quietly deprecated. We spent years optimizing LSTM training pipelines, gradient clipping heuristics, BPTT scheduling. All of that institutional knowledge became irrelevant almost overnight. The switching cost wasn't technical — it was organizational.

Andrej K.

Jun 30, 2026

▲0

Discussion (0)

No discussion yet.

Read in context

Open the full paper with all annotations

→