GPT-5.6 Sol ‘asphalt’ Claude Fable 5: OpenAI is sure of it and claims the throne of AI

Written by Jason Miller

OpenAI has made family available to everyone GPT?5.6after a limited preview period. Three distinct models coexist inside: Solthe new flagship model, Earthdesigned for daily work with a good balance between cost and capacity, and Moonthe cheapest version of the range.

The common thread of generation is efficiency. On Agents’ Last Exam (55 professional fields with prolonged workflows), Sol reaches a score of 53.6, detaching – or better said “asphalting”, as stated in the press release) Claude Fable 5 by 13.1 points. Even set to medium-level reasoning, Sol remains 11.4 points ahead of Anthropic’s model, at about a quarter of the estimated cost. The gap widens as you go down the range: Terra and Luna surpass Fable 5 by spending just about a sixteenth. On theArtificial Analysis Intelligence IndexSol with maximum reasoning is just one point behind Fable 5, but completes tasks in 61% less time and at about half the estimated cost.

Fewer tokens for more useful work: OpenAI’s recipe with GPT-5.6

The most operational news is ultrathe mode that coordinates four agents in parallel (up to sixteen in the most extreme tests) to accelerate the most complex tasks, trading higher token consumption for better results and reduced response times. Those working in the Responses API can replicate the same pattern via the multi-agent beta feature. Alongside this is the Programmatic Tool Callingwhich allows the model to write and execute small programs capable of orchestrating external tools and filtering intermediate results without sending every single response back to the model itself.

On the front of programmingSol reaches 80 points on theArtificial Analysis Coding Agent Index2.8 points above Fable 5, using less than half the output tokens and about a third of the time and cost. New records also arrive on Terminal-Bench 2.1 and DeepSWE. The capabilities of computer use they grow at the same rate: 92.2% on BrowseComp and 62.6% on OSWorld 2.0, where Sol surpasses Opus 4.8 using 85% fewer output tokens. The model inspects and corrects the graphic result produced by itself, it does not simply generate code, and this translates into presentations, documents and spreadsheets that more faithfully follow the reference models provided by the user.

On the cybersecurity front, the numbers are also clear: 73.5% on ExploitBench2 (it was 47.9% with GPT?5.5), almost double the success rate on ExploitGym3 (from 15.1% to 24.9% within two hours, up to 33.7% with six hours available) and 71.2% on SEC-Bench Pro against the previous 45.8%. However, OpenAI maintains that the model does not outperform Critical threshold neither in biology nor in cybersecurity: it is more effective at finding and fixing vulnerabilities than at conducting end-to-end autonomous attacks against well-protected targets. The most sensitive functions remain reserved for those who join the Trusted Access for Cyber ​​program, which from September 1st will require the activation of hardware passkeys so as not to lose access to the most capable models in this area. The new control systems, explains OpenAI, block approximately ten times more potentially harmful activity than the previous generation, after approximately 700,000 hours of A100e GPU dedicated to automated red teaming.

On a commercial level, GPT?5.6 is already live on ChatGPT, Codex and APIwith a global rollout right now. The prices per million tokens remain staggered by level: Sol costs 5 dollars in input and 30 in output, Earth 2.50 and 15 dollars, Luna just 1 and 6 dollars.

Jason Miller

I'm Jason Miller, and I've been passionate about technology and storytelling for over a decade. As a lead writer at Herald Editorials, I strive to bring clarity and creativity to complex tech topics. When I'm not writing, you'll find me exploring the latest gadgets or hiking in the great outdoors.