GPT-5 Turbo Benchmarks: What Actually Changed for Real Workflows
We ran GPT-5 Turbo against Claude 4.5 and Gemini 2 across 60 real tasks. Here are the honest results.
Every launch chart looks amazing. Real workflows don't. We tested the three frontier models on the tasks our readers actually do.
Methodology
- 60 tasks across writing, coding, analysis, research
- Each task run three times, blind-graded
- Ties broken by median completion time
What surprised us
GPT-5 Turbo won on breadth. Claude won on prose. Gemini won on cost per task at volume.
Keep reading
Hand-picked next steps based on what you just read.