Gemini 3.8 Flash launch with only 200k context window

Describe the Bug

Gemini 3.8 Flash launch with only 200k token window

Steps to Reproduce

Just send the request.

Expected Behavior

1M token window

Screenshots / Screen Recordings

Operating System

Windows 10/11

Version Information

Version: 3.19.7 (system setup)
VS Code Extension API: 1.128.0
Commit: 90de2327392570a5f5f625c656c6749d228e6430
Date: 2026-09-02T23:03:07.739Z
Layout: IDE
Build Type: Stable
Release Track: Nightly
Electron: 42.10.0
Chromium: 148.0.7778.280
Node.js: 24.18.1
V8: 14.8.178.38-electron.0
xterm.js: 6.1.0-beta.291
OS: Windows_NT x64 10.0.22631

For AI issues: which model did you use?

Gemini 3.8 Flash Medium

For AI issues: add Request ID with privacy disabled

09a71acc-8678-4b07-8cb6-0f2ca8bd3fef

Additional Information

Does this stop you from using Cursor

Yes - Cursor is unusable

Same with Gemini 3.7 Flash High 85d07d03-ddfd-4dfa-b0cb-8c9ef3211d79

You’ve probably broken all the Other models that don’t have a context window selector.

Hey @Artemonim, thanks for the report.

This is the same underlying limitation as your Kimi K3 thread, and I’ve added this report to the issue we’re tracking. Gemini 3.8 Flash currently has no Context option in the model picker, so every chat runs with the standard 200K working window (unless it has gotten stuck in the now-invisible Max Mode, which bumps it up to 1M).

In that case, both Gemini and Kimi become useless for long tasks. For example, right now I consider Kimi to be the ideal model for a pipeline where I might have 300,000–400,000 tokens at the end, and those tokens can’t be compressed because the pipeline breaks down.

OpenAI’s 1M models are slightly worse for these tasks than K3, and on top of that, they may become unavailable in November.