Grok 4.5 really tires me out

I’ve been working with Grok 4.5 almost exclusively for the last 7 to 10 days (I’ve lost track of time). I chose to do a long, complex project with it in order to really get a feel for it and evaluate it strongly. We are working from a master plan with 18 distinct phases and we create separate specialized plans for each phase 1 at a time. We are just past halfway through.

But this model seriously drains me. I have to constantly fight with it to do proper planning and debugging. And it guesses SO MUCH without first investigating and looking at existing working code for examples to solutions. We are constantly re-inventing the wheel over and over again. It makes the same mistakes over and over again even when the context hasn’t yet summarized. And it’s even deleted tests on me outright because it couldn’t get them to pass – so I had to fight really hard with it to get it to realize that it was NOT the tests but it’s implementation. Dang thing eventually figured it out, but omg it’s like fighting with a kindergartner. And it seems to think that when making a plan to fix an unforeseen problem is to write a plan to INVESTIGATE (i.e. don’t do ANY investigation until I “Build” the plan) instead of investigating first and then documenting findings and creating a plan with identified fixes. It’s preferences, by default at least, are just backwards.

My experience with models from Anthropic haven’t been this exhausting. But unfortunately I feel that I need to see the rest of this master plan through with it because its plans are written SO differently from plans from other models (i.e. they are barely human understandable – lots of short snippets with few details) I fear that other models aren’t going to understand what to do from Grok plans.

FWIW, I have tons of rules and skills – a very mature repository for AI agents to use.

For me auto mode has consistently proven to be the fastest, most accurate and certainly most economical.

Hey @RayGraham!

Thanks for sharing this feedback, and sorry your experience with Grok hasn’t been great.

Instruction following is a huge area of focus for us, and one where we know Grok has a lot of room to grow. I’d be curious to get your take on a few things:

  • Have you seen this behavior mostly in Cloud Agents, local agents, or both?
  • Are these longer chats? Grok’s context window currently maxes out at 256k.
  • If you can reproduce the behavior with Share Data enabled, grab the Request ID from that chat and share it here. Request IDs like these help the team build evaluations as we develop new models.