Hey, good question. I’ll break it down point by point.
Effort controls how much the model “thinks” before answering. Grok 4.7 and Grok 4.6 have four levels: Low, Medium, High (default), and Extra High. The higher the level, the more the model reasons through harder tasks.
On your second question, it’s both. The price per token is the same at all levels, but higher effort makes the model generate more reasoning tokens, and those count as output tokens. So a higher level uses more of your usage and takes noticeably longer for each step. Separately, there’s a Fast toggle, and that’s what changes the per-token rate. Effort does not.
How to pick based on the task:
Low: quick and straightforward edits, renames, small fixes, or questions about code you already understand.
Medium: most everyday tasks. A good balance of speed, cost, and quality.
High default: multi-file changes, tricky debugging, or tasks where you want the model to double-check itself.
Extra High: genuinely hard problems, big refactors, or when High didn’t work. Longest wait and highest usage.
For day-to-day programming, High or Medium is usually the best choice. If you’re watching usage, keep Medium as your default and bump a specific hard task to High or Extra High when needed. The level is set per chat in the model picker.