Grok Bot voice calls should integrate with the chat session (post turns + `three-part-summary`)

Feature request for product/service

Chat

Describe the request

Problem

Voice calls do not integrate with chat session turns. Spoken work stays on a separate rail (call JSON at hang-up; chat mostly gets a duration receipt). That is a big usability gap: the mode you use to talk is not the same record the bot and you rely on for context.

Why Dictate is not the answer

Composer Dictate (mic on the prompt field) still lands text in the chat, but it is keyboard-mode with a mic: start a draft, speak, stop, edit, send, then optionally jump into a call. Switching between Dictate and voice-call is clunky and not hands-off. A voice UX should prioritize staying on the call — compose, post into the chat session, hear the labeled reply — without touching the composer.

Framing (intentional separation of concerns)

If the voice-call lane is kept as its own transcript, treat that as an intentional separation of concerns—not an accident—and decorate anything that comes from the chat session so the division is obvious. Voice may compose and narrate; the chat session remains the durable turn record. Cross-lane content must be labeled as chat-session material whenever it is spoken on the call.

Request

Treat voice as a prompt-entry assistant for the existing chat session:

  1. Draft and refine prompts on the call (with trained knowledge and web fetch as aids).
  2. Simple in-call commands to post a composed turn into the chat so it becomes a normal session turn and enters that context window.
  3. Quick re-entry: “give me a three-part summary of the chat session” (three-part-summary):
    • Part 1 — one sentence: overall chat focus.
    • Part 2 — one sentence: roughly the last third of the history.
    • Part 3 — as many sentences as needed: detailed catch-up on the last handful of turns.

Explicit cross-lane speech (required)

  • three-part-summary is explicitly a summary of chat session history. The user already means that; the voice assistant must introduce the spoken result that way, e.g. “Here’s the chat session three-part-summary…” then the three bands—not a vague “here’s a summary” that could be the call itself.
  • After the user asks to post a turn into the chat session, the voice lane submits that prompt, waits for the chat-session agent’s reply, then speaks it with an explicit handoff, e.g. “Here’s the reply to your prompt from the chat session agent…” followed by that reply.

Why

Voice is hard to use as a first-class mode while it stays disconnected from the chat’s turns. If the lanes stay separate, the UX must make the boundary audible. Closing the gap with post + labeled summary/reply keeps one discourse record and makes returning to a long thread cheap.

Acceptance

  • three-part-summary is grounded in the bot’s chat session (not only the call log) and is introduced as a chat session three-part-summary.
  • Posted turns appear as normal chat messages; when the chat agent replies, voice speaks that reply after an explicit “from the chat session agent” lead-in.
  • Cross-lane content is always decorated so the user can tell chat-session context from call-only talk.

Naming note

No established product term matches this exact three-band shape (whole-thread focus → last-third → last-handful detail). Closest jargon elsewhere: nested / multi-resolution summary, pyramid summary, or zoom levels. three-part-summary is the command key for this FR.

Hey, thanks for the detailed feature request. It’s laid out really clearly.

A couple notes that might help right now: part of what you’re describing already works to some extent. If you ask the Bot to do something on a call, it passes the request to the chat agent and then speaks the finished result back on the call. So the basic handoff from call to chat is already there, even though the relay doesn’t show up yet as a normal visible turn in the chat.

On the main idea, that calls and chats currently live as two separate records, this is a direction we have on our radar. Separately, three-part-summary and explicit cross-lane labeling like “from the chat session agent…” are narrower and newer parts of the request. I’ve passed those to the team along with the rest.

I can’t share a timeline yet, but the request is logged. If there’s an update on this, I’ll reply here.