Deepseek V4 Flash 0731 - Model Request

Please add DeepSeek V4.1 Flash ASAP

Please consider adding DeepSeek V4.1 Flash to Cursor as soon as possible.

This is not just another cheap “Flash” model. It offers an unusually strong combination of speed, intelligence, coding/agent capability, and price — without the typical “fast = much less intelligent” trade-off.

Why V4.1 Flash?

Metric DeepSeek V4.1 Flash
Artificial Analysis Intelligence Index 40
Artificial Analysis output speed ~210 tokens/s
Time to First Token (TTFT) ~1.37s
Terminal-Bench 2.1 90.6
DeepSWE v1.1 74.2
Codeforces Rating 3471
Context Window 1M tokens
API Input $0.30 / 1M
API Output $1.20 / 1M
Off-peak Input $0.15 / 1M
Off-peak Output $0.60 / 1M

Source: Artificial Analysis + DeepSeek official API/benchmark data.

1. Extremely fast, without the usual intelligence downgrade

Artificial Analysis currently measures V4.1 Flash at approximately 210 output tokens/sec, while giving it an Intelligence Index score of 40.

DeepSeek itself states that V4.1 Flash has surpassed V4 Pro in performance, cost, speed, and total runtime.

This is exactly what a coding agent needs: Flash-level latency without paying a huge “intelligence tax.”

2. Very strong coding & agent performance

This isn’t just a fast chat model:

  • 90.6 — Terminal-Bench 2.1
  • 74.2 — DeepSWE v1.1
  • 65.4 — NL2Repo-Bench
  • 3471 — Codeforces Rating
  • 90.9 — GPQA Diamond

These are particularly relevant to Cursor because repository understanding, terminal interaction, tool use, and multi-step software engineering are exactly where an agent model needs to perform.

3. The price/performance is extremely compelling

Model Input / 1M Output / 1M
DeepSeek V4.1 Flash (off-peak) $0.15 $0.60
DeepSeek V4.1 Flash (peak) $0.30 $1.20
Composer 2.5 $0.50 $2.50
Grok 4.6 $2.00 $6.00
Composer 2.5 Fast $3.00 $15.00

At off-peak pricing, V4.1 Flash output is roughly:

  • 10× cheaper than Grok 4.6
  • 4.2× cheaper than Composer 2.5
  • 25× cheaper than Composer 2.5 Fast

For agentic coding involving hundreds of model/tool calls and repeated edits, this difference is enormous.

4. Developer adoption is already happening

The ecosystem is already moving quickly.

OpenCode is an official V4.1 Flash launch partner, and its public usage data currently reports approximately 281K unique users and 4.59M completed sessions for V4.1 Flash, ranking it #1 in recent OpenCode usage.

Other coding-agent ecosystems, including CommandCode, are also moving quickly around low-cost, high-performance models like V4.1 Flash.

This demonstrates real developer demand for exactly this combination: fast + intelligent + cheap.

Cursor should not be late to this model.

Request

Please add DeepSeek V4.1 Flash to Cursor’s model selector as soon as possible.

For high-frequency coding and agent workflows, it has the potential to be one of the best price/performance models available today — and importantly, its speed does not appear to come from simply making the model dramatically less capable.