Kimi k3 disoriented and expensive in cursor

Where does the bug appear (feature/product)?

Cursor IDE

Describe the Bug

in about 5 minutes, kimi k3 burned ~$20 of on-demand api use by reading sequential ~40 LOC blocks of a single file in a disoriented state. Had to be stopped.

It was asked to overhaul a set of UI elements which was not overly challenging

Steps to Reproduce

asked kimi k3 to overhaul a set of ui elements, it got disorientated (expensive)

Expected Behavior

would of expected it to not:

  • get disorientated and loop through a file for no apparent reason
  • burn through on-demand use on with zero gain

Screenshots / Screen Recordings

Operating System

Linux

Version Information

Cursor 3.13.25

For AI issues: which model did you use?

kimi k3

For AI issues: add Request ID with privacy disabled

sorry, don’t have this

Does this stop you from using Cursor

No - Cursor works, but with this issue

Kimi K3 from Cursor is very low quality! Apparently very quantized.

I had similar problems right now, it went on crazy loops debugging an app just repeating the same script over and over, dumber than any small sized coder model.

This behavior simply does not match native Kimi K3 from other CLIs

Is this on purpose so we are forced to use Grok (new rule from SpaceX)?

Cursor is great, but if it continues to do things like this, we will leave

I am just discovering that it is not worth it to use Kimi and GLM models via Cursor, it is better to use their native harness instead.

Direct providers means non-quantized models, and way cheaper usage.

It is kinda sad, because Cursor is a good terminal tool

Hey @Josh_Barnett, thanks for the report and the screenshot!

In short: the model got itself stuck in a repetitive loop, re-reading your file in small sequential chunks without making real progress, and our automatic loop protection didn’t catch it because each step looked just different enough to slip past. I’ve shared these findings with the team.

Nothing here is intentional, we don’t intentionally degrade model quality. This was the model misbehaving on this session.

Thanks for reply @Colin

It did seem disproportionately expensive - felt like I would have struggled to use that many tokens if i tried (i.e. Fable max).

Any chance of a token credit?

Any questions about billing/credits need to go through [email protected]!