Hey, thanks for coming back with details, and I’m glad Grok 4.7 is already a bit better.
For quality, the strongest signal for the team is a thumbs down on the exact runs where the model went off track. That way, that specific run gets queued for review. Also, your notes in the Grok thread are really on point: Share your Thoughts on Grok 4.6. Specific examples like the i18n case with inline defaultValue are exactly what helps.
A couple things that really affect results right now:
- Add a Project rule for your i18n convention (where keys live, no inline
defaultValue). That keeps any model disciplined, not just Grok. - Check the effort level and the Fast toggle in the model picker. Cloud Agents often default to the Fast variant, and that can feel like “40 minutes with no progress”. Explicitly setting non-Fast plus the right effort usually makes a noticeable difference.
About “we reread every line and don’t merge an MR without edits”, I get it. It’s exhausting, especially if trust used to be higher. Models are constantly being retrained, and this is an area the team is actively working on. I can’t give an ETA, but feedback via thumbs down and the thread directly affects priorities.
If you hit a specific run where things go really badly, send the link (cursor.com/agents/…) and I’ll take a look at that run separately.