H

Horace H.

3 annotations0 followers0 following

FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness

"even with the increased FLOPs due to recomputation, our algorithm both runs faster and uses less memory"

This is the result that broke my intuition about GPU optimization. Recomputing the forward pass costs FLOPs; saving and loading from HBM costs time. On H100s, the ratio is roughly 3:1 — memory is 3x more expensive than compute per unit of work done. FlashAttention exploits this ratio; most ML researchers hadn't internalized it.

▲ 0discuss →

Mixtral of Experts

"only uses 13B active parameters during inference...outperforms Llama 2 70B and GPT-3.5"

From a systems perspective: 13B active params but 47B total means you still pay the memory bandwidth cost of loading 47B params, you just compute with fewer of them. The speedup is real but comes from FLOP reduction, not memory reduction. For memory-bound inference (which most LLM serving is), MoE's efficiency story is more complicated than the benchmark numbers imply.

▲ 0discuss →

FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness

"a missing principle is making attention algorithms IO-aware"

The IO-awareness framing changed how I think about GPU programming entirely. Before FlashAttention, I'd profile FLOPs. After, I profile memory bandwidth. These are not the same thing on any modern GPU — and the fact that it took a 2022 paper to make this obvious to the ML community reflects how thoroughly we'd been taught to think in compute abstractions rather than hardware reality.

▲ 0discuss →