On demand costs much higher?

I used like 90% Grok 4.5 High and took me about 2 1/2 weeks (daily work) to fill up the included Pro+ $60/mo Cursor models bar. Then it took me about 2 days to fill up the left over 50% to 100% in the ‘Other Models’ bar. Now today It takes me less then 4 hours to spend $32 for using Grok 4.5 High (and now also Medium).

Can someone explain these differences in costs?

Hi @Someguy , the ‘Cursor Models’ and ‘Other Models’ buckets are billed at different rates. @mohitjain provided an excellent explanation of this recently in this thread: Unclear usage billing - #10 by mohitjain

Thanks for your pointer. However I don’t use subagents, they are blocked on purpose (also for this reason) plus I never used the medium grok, always high during the included cursor models period. Or am I not understanding this?

The quota for Cursor models is five times larger than the quota for third-party models. If you’ve used up your quota, you’ll switch to On-Demand.

When used properly, subagents can improve the quality of work and reduce costs.

Thanks, but does that also mean that the Cursor models (Grok/Composer) are also x5 when billed in On-Demand?

No and you can check via https://cursortokens.vercel.app/ or https://cursor.com/dashboard/billing

Hi all! Glad to see you all helping each other out here in this post. I’ll add my 2 cents although it’s not really neeeded.

Once the Cursor Models allowance is exhausted, Cursor Grok 4.5 and Composer 2.5 can continue against the smaller Other Models allowance. It can feel like the other models pool is moving faster even though the model’s token rate has not changed. (The Cursor Models pool has more included usage than the Other Models pool).

After both included pools are exhausted, on-demand usage is billed at the same published per-token rates, with no additional multiplier or markup. High versus Medium does not change the listed per-token rate, although effort can affect how many tokens are used.