“passes a simulated bar exam with a score around the top 10% of test takers”
The 'simulated' qualifier is doing a lot of work here. It's the multiple-choice MBE portion, not the full bar. Real bar exams include essays and performance tests that require sustained legal reasoning — a very different capability.
Law student
Jun 29, 2026
Discussion (0)
No discussion yet.
Read in context
Open the full paper with all annotations
More annotations on this paper
“GPT-3.5's score was around the bottom 10%”
This comparison is the real story. The jump from 3.5 → 4 on legal reasoning is massive. What's wild is GPT-3.5 was already considered impressive when it launched — and it was apparently failing the bar at near-chance level.
“GPT-4 exhibits human-level performance on various professional and academic benchmarks, including passing a simulated bar exam with a score around the top 10%”
The bar exam result is real but the framing is misleading. GPT-4 sat the multiple-choice MBE portion under conditions that don't replicate actual exam constraints — no time pressure, no physical fatigue, no adversarial essay scoring by human graders. Passing a simulated exam and passing the actual bar exam have different success rates for the same reasons that in-context benchmarks don't transfer to deployment. 'Top 10%' is a specific number attached to an imperfect simulation.
“we developed infrastructure and optimization methods that have very predictable behavior across multiple scales. These improvements allowed us to reliably predict some aspects of the performance of GPT-4 from smaller models trained using 1,000× – 10,000× less compute”
The predictable scaling claim is the most scientifically significant sentence in the report, and also the least scrutinized. OpenAI doesn't disclose which capabilities were predicted, what the prediction error was, or whether the method works for emergent capabilities (which are definitionally hard to predict). 'Some aspects' is doing enormous hedging work. If scaling predictions were fully reliable, the field would have a much cleaner path to forecasting dangerous capabilities — which makes the imprecision here consequential.
“the post-training process, the calibration is reduced. The post-training hurts calibration significantly”
This admission — that RLHF makes the model more confident and less calibrated — is buried in Limitations and rarely cited in discussions of GPT-4's reliability. A well-calibrated model's stated confidence tracks its actual accuracy; post-training breaks this correspondence. The model becomes more useful (follows instructions, sounds confident) and less trustworthy (its expressed uncertainty no longer reflects reality) simultaneously.