The current pricing says that Pro+ should have 675 sonnet 4 requests, I ran less than 70 requests before hitting my limit. I switched to Ultra, and after half a day of using it, the API cost was $35 and it used 65 million tokens (58 million from cache read) in half a day (when my input tokens were only 564,901 and output token 264,911! Other customers have reported 93 million cache read tokens were used for a couple thousand input tokens (120x more than it should be). People on PC and Mac have confirmed this, specifically with Claude 4 Sonnet and Opus.
The issue is very likely that the Cache Read token usage spiked up to 120x usual rates (See images for proof), this is why people are using up their credits in days when it used to last all month. I think this is the main issue cashing most push back, where people are getting billed high amounts. Its not the new pricing plan, its this bug that is charging 70-120x more and draining peoples account credits.
I’ve literally used Cursor every day for the past year +, developed 12+ applications with it. I’ve used the pro, then metered usage pricing up until last month (for 7+ months with metered usage at same rate of use it would be $60/month in total usage every month this year). I made 10 minor updates to my webapp today, that I’ve worked on with cursor for 5+ months without this issue. With this issue, with the same amount of use, I’m getting charged $35 for a half a days usage! Look at my image, it literally processed 63 MILLION TOKENS! in half day when my Input tokens was only 564,901 and output token 264,911 This is a big issue.
It will use up to 120x more Cache Read than input and output tokens combined! This is a huge bug that is costing people hundreds to thousands of dollars.
Operating System
Windows 10/11
Current Cursor Version (Menu → About Cursor → Copy)
I’m using CC as well now, which is good, but better for larger sets of tasks than quick vibe code updating like cursor. I would glady pay for the cursor Ultra $200/month plan if it wasn’t charging for 120x more tokens than I use. It’s such a bleed to spend $35 of the ultra credits in half a day. It would could me thousands to use it everyday of the month like I have for the past year. They need to fix this so our usage matches their claims for the new pricing plans, which is fine by me without the cache write bleed:
“Expected usage within limits
Expected usage within limits for the median user per month:
Pro: ~225 Sonnet 4 requests, ~550 Gemini requests, or ~650 GPT 4.1 requests
Pro+: ~675 Sonnet 4 requests, ~1,650 Gemini requests, or ~1,950 GPT 4.1 requests
Ultra: ~4,500 Sonnet 4 requests, ~11,000 Gemini requests, or ~13,000 GPT 4.1 requests”
I got maybe 70 sonnet 4 requests on pro+ before hitting my limit in 7 days, when it used to last all month without ever running out for months on end with the same project.
If this is a bug I want my tokens back. If this is working as it should, I’m gone, this is beyond expensive. The fact we do not even get a response from the team is destroying all trust
Based on my personal testing, it’s clear that there are no issues with token usage in Ask Mode, and the problem seems to lie in Agent Mode.
It appears that a significant number of tokens are being consumed during code modifications or similar processes in Agent Mode.
In particular, it seems that a large number of tokens are being consumed during the process of modifying code.
@sant_dev yes this is the same issue we are having, probably everyone using Claude 4 Sonnet/Opus. I drained through my entire Pro+ account in 7 days as a result, switched to ultra and drained $35 in half a days use. It’s like burning money. I’m sticking with CC until this is fixed.
Can’t agree more. You wrote it perfectly. It’s so sad that people put their trust in a company, like myself buying the ultra plan, and then being told your the problem by cursor ambassadors on this forum. I legit went on a fresh project with nothing on it, no chats, no files or folders, and I said test. 9850 tokens used, and I supposed to believe it’s on me. I already switched to Kiro and am most likely going to chargeback if they don’t give me a refund.
Cursor admitted today that they are baiting and switching annual users who paid for a amount of requests per month model, as I showed you the pricing site still shows a you get x requests per plan model. However cursor said it no longer a per request model it is a straight API fee model. Which is the bait and switch I was talking about earlier, customers who signed up annually based on processing by requests are being switched out of their contracts. This is exactly what I described, Cursor no longer holding to their advertised request usage promises.
Hey, we are currently investigating the reports of higher-than-expected cache read tokens, and if an issue is found, we will make things right with anyone affected!
Cursor’s system prompt should take ~5000 tokens on it’s own in a new chat, which should be cache reads only (the cheapest type of token) so seeing such high token usage does seem abnormal.
To clarify your last statement (@Illuminationx), you can see the exact details of our pricing here, but the numbers around how many requests can be done is the average across all Cursor users based on the standard length of a request. Some users may use more or less depending on what their input is to the LLM.
For those who signed up to an annual contract, they could either opt out via cursor.com or still can opt out (annual only) via [email protected] and they will remain on the old structure until the end of their billing period.
As we look into this issue, we have so far been unable to reproduce an issue internally to cause a drastically high cache read token amount.
Is anyone able to, with a brand new chat, trigger this reliably?
If so, if you can disable privacy mode and send over a Request ID using this guide, that would really help the team look into this, and fix anything that may be broken!
Thank you, we really appreciate looking into it depper. I have used cursor every day for over a year, it was my #1 go to, if this isn’t resolved I’ve already started using Kiro.dev and won’t look back. 100s of other people are having this same issue, I’ll share some of the screenshots for clarity that this is a wider spread issue.
My pro+ and metered usage never exceeded $60 on the 1 project I’m working on, that is same size for 7+ months, as of this update my Pro+ tokens ran out in 7 days. I’ve never not made it to the end of the month on the same project and same amount of work, people are reporting tokens draining in 3 days on pro now, when it used to last all month for them.
I updated to Ultra $200/month and within half a days use I was at $35 of usage, this is unruly, a 10x+ drain on tokens vs all previous months.
Here is some of the proof that hundreds of users are having this same issue of massive token drain. I won’t post all 100+ screenshots but here is some:
Also it’s important to mention people only seem to mentioning this issue with Claude 4 Sonnet & Opus, not the other AI models. When using Sonnet on Ask mode, it doesn’t generate the large Read Cache.
Not only get a refund for the overage of tokens being used. But also my time back that I lost not being able to work because of this issue. I lost 5-7 days waiting for this to be fixed, and I don’t want to lose that time.