The last 36 hours have been incredible. It’s really great to get a model that’s slightly better than GPT-5.5 for the price of GPT-5.4, which runs a model at the level of GPT-5.4 for the price of Composer 2 Fast.
oftop
I think after half a year of thinking about it, I’ll finally pull myself together and update my Agent Compass to celebrate this teamwork
Incredible 24 hours.
Grok 4.5 nipping at Opus 4.8’s heels
GPT 5.6 sol nipping at Fable 5’s heels.
Great to see some competition at the top end again and an alternative to Fable’s prices.
I’ve noticed the prompts it generates for sub-agents are a bit too generic. I tried the same prompt against Opus, Fable, Composer 2.5 and GLM 5.2 and they give a much more focussed prompts to sub-agents, so the results end up better.
Is this something that can be tweaked by the Cursor team with their system prompt?
Example request ids, same prompt against different models. Look at the first ‘exploring code-base’ sub-agent prompts:
Great experience in the first 24 hours. Cheaper than Opus 4.8, at least on the same level if not higher (subjective impression). Not seeing big differences in output between GPT 5.5, Opus 4.8, GPT 5.6 though, all of them are quite perfect already (with proper workflows and context engineering of course). The major difference so far had been that Opus 4.8 was better at visual design than GPT 5.5, though none of theme were particularly good at it, so adding visual design skills (e.g the “taste” skill) was still needed. I’m excited to see whether GPT 5.6 has better visual taste, as they promised
Luna-Max does it pretty. It also has a slightly better GDPVal than Grok 4.5, so I choose it as the leader and Grok as the main SWE, with the possibility of escalating to GPT-5.6-Sol-High.
I discussed the benchmarks and X.com feedback with Grok and Fable, where people generally seemed to reach a consensus that, in terms of price-to-performance, the best choices are either Luna-Max, Sol-Medium, or Sol-High+.
That said, I still have a feeling that Terra-XHigh would be the better option for tasks involving large context windows. Fable actually thinks it’s better to go with Terra-Max if you’re looking for a cost-efficient premium.
Fun fact: Subagent Senior Terra called Subagent Verifier Pro herself before returning the report to the team leader.
Plan by GPT-5.6 Sol + Grok 4.5 (before that team call)
Leader GPT-5.6 Luna Max 1M
Grok-4.5 + (GPT-5.6 Sol High, GPT-5.6 Terra Max) as SWE
Composer 2.5 + GPT-5.6 Terra XHigh 1M as Verifiers
Worked for 4.5 hours
Honestly, I’d be curious to see how my multi-agent orchestration setup would score on a proper benchmark. Unfortunately, benchmarking would be pretty expensive…
model: gpt-5.6-terra[context=272k,reasoning=max,fast=false]
Subagent RID: No request ID found; called in 11:32:27 UTC+0
Leader RID: dec191af-427e-46d5-8d43-6c94f6864e01 (but this is follow-up, not the teamwork)