Personal assessment on cursor models

I have been developing large-scale software projects using Cursor for nearly a year. I heavily rely on auto mode, given that it offers affordable pricing. Judging based on my recent experience as the codebase grows, I would say that auto mode offers competent analysis capability. However, its coding capability often shows juniority, lets down what its analysis and specification decided on, results in half-cooked implementations, and breaches the clearly laid out specification. I think there is a gap in the auto mode in intelligently switching between the analysis (or specification) models and the coding models based on the complexity of the analysis (or specification). It costs me significant code review effort, and costs significant tokens and time to frequently refactor source code. I think I won’t be the only one here to expect better auto mode coding results upon each Cursor update.

In general, I agree with this. I pretty much only use Grok 4.6. Composer 2.5 was good for awhile, but I found (as you said) it was more of a junior coder that needed small batched tasks with clear instructions. Grok 4.6 is like a senior developer, but is not an expert (like Opus Max).

I never use auto because Model selection is still so important.

Thanks for suggesting Grok4.6. I am going to give it a try. But the advantage of Cursor is its auto mode. It is losing its advantage when its coding outcome is at a low quality.