Article
Two paths in one night: Opus5.5 surged to first place, while GPT-6 cut costs in the same tier by half
Dongcha Beating AI Brief: Artificial Analysis has completed third-party testing of several new models released last night. Claude Opus 5.5 max scored 58 points on Intelligence Index v4.3, currently ranking first. GPT-6 Astra and Fable 5.1 both scored 53 points, while Opus 5 scored 51 points. Opus 5.5 achieved the highest score in 10 of the 6 evaluations.
However, this top score also used more tokens. Opus 5.5 max generated approximately 119000 tokens per task on average, about 60% more than the 73000 generated by Opus 5. With API price cuts of 20% and cache-read price cuts of 60%, the final cost per task was 5.98 USD, essentially equal to the 5.86 USD cost of Opus 5. The more practical medium tier scored 51, matching Opus 5 max, while costing only 1.34 USD per task.
OpenAI has taken a different path. GPT-6 Sol and Luna performed similarly to the previous generation of GPT-5.6 on the Intelligence Index, but their prices fell significantly. Sol max's per-task cost dropped from 1.99 USD to 1.06 USD, while Luna's fell from 0.18 USD to 0.07 USD. In the coding Agent evaluation, Sol rose from 55 points to 57 points, with costs also falling by approximately half; Luna fell from 43 points to 41 points, but its per-task cost was approximately 60% lower.
The two companies' approaches are beginning to diverge: Opus 5.5 is using more reasoning to push the capability ceiling to a new high, while GPT-6 Sol and Luna are primarily making existing capabilities cheaper. On Artificial Analysis's performance/cost chart, the new models from both companies occupy new Pareto frontiers.
Original link: https://m.theblockbeats.info/flash/368584