GPT-4o System Card
"matches GPT-4 Turbo performance on text in English and code"
This claim is heavily qualified — 'matches' on benchmarks doesn't mean equivalent in practice. In coding evals GPT-4o actually regressed on some HumanEval subsets. Read the appendix before taking this at face value.