Voice input disappears after attaching an image

Voice input disappears after attaching an image

Currently, I can use voice input normally when the composer is empty. However, as soon as I attach an image/screenshot, the microphone button disappears and is replaced by the Send button.

This means I have to use voice first, finish the transcription, and only then attach my screenshot.

It would be much more natural to support:

Attach image → use voice to explain the image → send

Screenshot + voice is particularly useful when working with UI/code issues, since I often want to attach what I’m seeing and then verbally explain what I want changed.

Ideally, attaching an image shouldn’t disable or hide voice input.

Hey, thanks for the feature request. The logic makes sense: attach an image → explain with voice → send. For UI and code tasks, that really is more convenient.

This has already been discussed here: Speech-to-Text Functionality Issues with Image Pasting in Agent Mode. I’d recommend checking it out and adding your thoughts or details there so we can keep the signal in one place.

I can’t share an ETA yet, but I’ll post an update if we have news on this.