Absolutely unusable token usage

I’ve been on the Ultra plan since February after being on the Pro plan for around a year before that (perhaps longer?), and the usage calculations lately have become so absurd that continuing to pay for this service just seems crazy.

Back around March or April, I was able to have GPT-5.5 xhigh work though forty minute agentic sessions on a complex codebase, reading and writing thousands of lines of code, at a cost of maybe $15-$20 worth of tokens (~5% of my monthly quota). Even with a 1M context window, I could work most of the day and hardly exceed 10% of my quota. I actually found it difficult to use the entire monthly quota if working efficiently.

Lately, if I have 5.6 Sol high (claimed to be more cost efficient) work for 10 minutes on a simple issue involving unit tests, (5 short prompts, not much work), it ends up costing $20 worth of tokens. A moderately complex task which requires 40 minutes of work will burn $100-150 worth of tokens (~25-30% of my monthly quota). With a single day worth of work, I can realistically expect to exhaust my entire $400 monthly quota of API credits. I can’t imagine using the Pro plan. I would exhaust my month’s usage after a single 30 minute session of fixing unit tests, or after a single prompt to implement some simple feature. It doesn’t make sense.

I love using Cursor, which I think is a great product. But paying $200 per month for Cursor to solve three moderate difficulty problems is just an unacceptable ROI.

(I’ve tried extensively to make Grok 4.5 work in Cursor, but it simply isn’t smart enough for the sort of targeted development in complex codebases that one actually requires Cursor for. Ask it any sort of pointed, technical question about a specific section of code, and there’s a 30% chance that what it says is untrue, applies to a slightly different question than the one which was asked, ignores conditions on the question or invents ones which do not exist, etc. 70% accuracy may as well be 0%, because nothing it says can be trusted. Similarly, for writing code, Grok is almost guaranteed to violate the scope of whatever it is it was requested to do, exhibit poor judgement given any moderately complex architectural decision, etc., again making it useless for all but trivial tasks. I really do hope the next iteration of Grok is a step-change improvement, because this version is an unconstrained slop cannon and cannot be trusted anywhere near a production codebase. Sorry for the brutal assessment, but the fact that Grok is cheap doesn’t make Cursor useful for me.)

Hi @tylercasper, Thank you for the post, and we appreciate you as a Cursor user, and we appreciate you sharing your honest feedback.

We are working every day to make our Cursor models smarter, more tailored to your needs, and have them anticipate your needs better (e.g. do a better job interpreting what your true goal is). Cursor Grok 4.5 is an incredible model. It’s the best one we’ve shipped by far. It is my personal favorite. However, we’re not done - this is not the final state of Cursor models. We recognize there’s a lot more we can do to improve our models, and we’re working towards that every day.

Some thoughts on the GPT5.5 vs 5.6 models: While the per-token prices for GPT-5.6 Sol are the same as GPT-5.5’s, Sol at high effort tends to spend far more reasoning tokens per task, and long agentic sessions compound that, which matches the cost growth you’re seeing. My hunch (and I think other people have echoed this online) is that these frontier models, like Sol and Fable, at lower effort rates can still deliver really good results that are a bit more price-competitive than their prior versions at Xhigh reasoning levels. Unless you have a really strong need to use Max or Xhigh reasoning, I would probably steer clear of those for cost reasons.