Mistral 7B
"language models may compress knowledge more than what was previously thought"
This throwaway conclusion hints at something the paper doesn't explore: what exactly is being compressed? It's not raw facts — Mistral 7B still hallucinates. It's more like 'how to manipulate language in ways that match factual distributions.' The compression is syntactic-statistical, not semantic.
Are Emergent Abilities of Large Language Models a Mirage?
"nonlinear or discontinuous metrics produce apparent emergent abilities"
This paper made a lot of people uncomfortable in a good way. The hardest follow-up question: are there tasks where we'd expect genuine phase transitions in capability? Theory says yes — tasks with sharp computational thresholds (like parity functions). But none of those have shown up in the BIG-Bench data. That absence is itself a finding.
GPT-4 Technical Report
"we developed infrastructure and optimization methods that have very predictable behavior across multiple scales"
The unreported part of this story is how many times the prediction failed and they retrained. Predictable scaling sounds like a solved problem in retrospect; building the tooling to achieve it iteratively is a different kind of research contribution than the paper suggests.