Not sure what’s going on with Grok 4.5 but it has blown through my usage faster than any other model I’ve used. I’ve been able to use Codex 5.3 far more without it filling up my usage.
Add a Subagent “Middle SWE”-like card, describe it as a main Subagent and prohibit the Agent from specifying model when calling.
Insanely expensive for what it delivers–and it’s still in the 50% off period??
Disabled. Won’t use it. If composer 2.5 is discontinued or prices are jacked up to force people into other models, it’s a deal breaker.
Hi Kevin.
Can you please shed some light on how is it that all the other labs don’t delay their models for EU?
Also, with Elon promising new Grok every month, those “few weeks” of a delay mean at least a half if not full minor release.
Will every new Grok and Composer model from now on be one minor version late for EU?
Team - tried many times but the AskQuestion tool seems to fail with Grok 4.5.
It tries and say it’s not in its native tool list and cant find the tool. As soon as I switch to composer 2.5 in the exact same chat it’s able to use it
if i choose “Auto” Model that mean the agent will use composer 2.5 or grok-4.5 to excute?
I have also encountered this problem
![]()
The model is definitely faster than Composer 2.5, and I really like that.
My biggest issue is that it doesn’t seem to understand what I’m actually asking very well. I often have to make my prompts extremely explicit before it gets them right. Even then, it tends to complete the task literally instead of doing it well. It doesn’t seem to consider edge cases, existing code, or the broader intent behind the request.
Another thing I noticed is that the plans and Ask mode responses are much less useful than Composer 2.5. They often feel pretty shallow and skip over important details, while Composer 2.5 usually gives me enough context to understand and trust the approach.
Also, I don’t know if this is only a Chinese issue, but the Chinese responses are surprisingly unpleasant to interact with. It’s hard to explain, but the tone often feels oddly irritating rather than collaborative, especially when the model has misunderstood the request. It doesn’t feel like you’re working with a helpful coding partner.
For macOS app development, I’m honestly not even sure which model I’d choose. The actual error rate doesn’t seem dramatically different in my experience. I was expecting Grok 4.5 to be a clear upgrade over Composer 2.5, but so far it hasn’t really met that expectation.
Agent is the best mode for any job.
Try adding /Verifier Subagent card powered by Composer 2.5 and a hook that will require Verifier to run at the end of the job. Perhaps this will soften the effect.
My workflow this morning: Grok 4.5 for planning → Composer 2.5 for implementing plan → Grok 4.5 to smooth out problems (failing tests, UI tweaks / “creative” changes). Everything seems to be great so far and it has barely touched my usage, to the point that I’ve double-checked on the website that Cursor was capturing things accurately (seems to be). Of course, I am usually careful to use the /caveman skill to try to reduce token usage.
For some larger scale changes, I had Opus 4.8 (via our enterprise Claude) draft longer documents which I then fed to Grok 4.5 to break into plans for Composer 2.5 to execute: again, everything worked well. This new model seems to occupy a nice space in the value/capability continuum: it’s definitely more capable than Composer 2.5. While I appreciate the affordability of Composer 2.5, it often gets fixated when trying to hunt down bugs and is unlikely to recover from a bad path once its made a decision. Grok 4.5, on the other hand, was able to 1-shot a couple of bugs that I initially let Composer try to fix (unsuccessfully).
Overall, this morning’s experience, while limited, has been a very promising start. It has a couple of small “bugs” (more like ‘not yet implemented’ things I’ve become used to with Cursor), most notably, when in a grilling session, Grok 4.5 did not present me with clickable options and instead just drew a basic table: not a problem at all, obviously, but little things like that do make this feel a bit less polished compared to other models at the moment (I’m sure this will be fixed quickly).
Sorry for two posts in a row, but after working more deeply with Grok 4.5 last night and today, I’m sold. The performance is amazing for the cost. It seems to have access to more interesting tools (maybe I’m wrong about this?): it recently (correctly) traversed an open api using cURL and wrote appropriate connectors in our app. It basically 1-shotted this after an in-depth grilling session. Cursor has moved the baseline from Apprentice to Journeyman.