Share your Thoughts on Grok 4.6

Main announcement · Blog


Now that Grok 4.6 is available, we’d love to hear how it’s working for you.

Some things we’re especially curious about:

  • How does it handle long-running, multistep work? Grok 4.6 is built to stay with complex tasks across many steps, including more self-testing and verification before moving on. Does that match your experience?

  • How is it on visual and interactive projects? We found it especially strong at turning a broad product idea into a working first version, with stronger first passes than Grok 4.5. If you’ve pointed it at UI, apps, or other interactive work, we want to hear how it went.

  • How does it compare to Grok 4.5, Composer, and other frontier models in your day-to-day? Anything noticeably better (or noticeably worse)?

The kind of long-horizon, tool-heavy work it’s built for is exactly what’s hardest to capture in a benchmark, so your feedback shapes where the model goes next.

Drop your thoughts below (and leave a Request ID if you have one)! We read everything. :folded_hands:

Is Grok 4.6’s context window still 256K?

Hi @Isia_Wang Thanks for the question! Yes, that’s correct, 256k context window. Btw, Grok 4.6 adds an XHigh reasoning capability (in addition to Low, Medium, and High effort options).

Btw, why are benchmarks performed on High instead of XHigh?

Hi @Artemonim Good to see you here, friend

The launch numbers are at High because that is the default effort, and it keeps the comparison consistent with the other models and third-party evals in that table.

You can still see the Extra High benchmarks here on the Evals On CursorBench: Cursor · CursorBench

Still seems to have the laziness issue (i.e. ‘finish from here’ or ‘do the rest’):
3fa45f12-5834-4cd8-87ba-0f85d6006e5f


I asked it specifically to write a query that groups by unit count but it just told me i should do it lol

Can you try reporting this with the Thumbs Up / Thumbs Down if you can access it in the Agents View?

Wait Grok 4.6 uses the other models budget? not the Cursor models? I’ll probably end up not using it until the end of the 2x usage window…

Also 4.5 & 4.6 show the effort level rather than the name

Hi @luchillo17 , Grok 4.6 pulls from the Cursor Models included usage. We merged a fix to clarify that text, that currently shows just Grok 4.5 and Composer 2.5. The fix is live in the web but it won’t be live in the Cursor app until the next production release version of the Cursor app (in the next few days most likely). Here’s how it looks on the web now:

Yes, this is intentional.

That’s part of it, but I think a bugbot review was made as subagent and it slipped my mind, likely why the 1% on other models budget, sorry.

Ah, yes, subagents can be spun up with other models - currently, the expected agent behavior is that the agent determines the best matching subagent model to use for the task, which can be different from the main agent model. We are exploring ways to add additional user control here, but right now, I don’t have an update to share on this.

I am using the grok4.6 model which does not consume the usage quota of my other models。
this is my cursor version info:
Version: 3.15.19 (system setup)
VS Code Extension API: 1.128.0
Commit: de07bee81cefe43461ebf4f40c3d2d78d15052a0
Date: 2026-08-11T05:22:54.627Z
Layout: IDE
Build Type: Stable
Release Track: Default
Electron: 40.10.3
Chromium: 144.0.7559.236
Node.js: 24.15.0
V8: 14.4.258.32-electron.0
xterm.js: 6.1.0-beta.291
OS: Windows_NT x64 10.0.19045

I‘ve been using Grok 4.6 xhigh all day, and I must say it’s quite good. But I found that its context window fills up quickly, is it caused by the reasoning effort or by the model itself?

Nope, not amused, dumber than 4.5
Trying to fix bugs it created already in 8 different seperate chats, switching back to 4.5!

Extremely slow compared to 4.5. Very simple tasks take long minutes. also not that smarter than 4.5 from what i’m seeing, and i’ve been using 4.5 all day long for about a month.

Now have been trying 4.6 for a few hours straight and it is just slow. calling many tools and just thinking endlessly and i’m “stuck” waiting for the output to be shown. even on “fast” mode

hi @zheng_gu Thank you for the post! This is correct behavior, Grok 4.6 draws from Cursor Models pool, the same as Grok 4.5 and Composer 2.5 do. We are updating the language in the app so that the usage indicator reflects that Cursor models include all Grok models:

Hi @Congzhi Thanks for the question! Yes, the context window is still 256k, same as Grok 4.5, and extra high does fill that window faster. Higher reasoning effort produces more thinking on each step, and those traces stay in the chat so the model can keep using them later. Extra High also tends to do more verification and tool use, which adds more conversation content on top.

If you want longer sessions in the same chat, maybe try switching back to high? Starting a new chat when you switch tasks also helps.