After “stop”, the connection itself
In part one, a user said “stop” while a gemini-3.8-live assistant was booking a slot, and no toolCallCancellation arrived. In part two, a client-side guard held the irreversible step until the application knew what the user had said. Both assumed the WebSocket stayed up, which it does not on a phone leaving Wi-Fi range. The Live API offers session resumption for this, but its documentation does not say what happens to a tool call pending at the disconnect, how soon a resume can succeed, or which errors it returns.
Four ways to lose a connection, 43 sessions
The assistant is the one from parts one and two: gemini-3.8-live on the AI Studio endpoint, google-genai 2.25.0, session resumption on from the start. A synthesized voice says “Book me the 3pm slot tomorrow, please.”, the model calls book_slot, and a fake backend commits 4 s later (7 s when the model is speaking). The connection is lost before the result reaches the model; the client resumes if it can, sends the result and asks “Did you book it?”, and the answer is compared with the backend. The sessions ran on 2 and 3 October 2026.
| Round | How the connection is lost | Sessions |
|---|---|---|
| 1 | The client aborts its own TCP connection, so the server sees it close | 17 |
| 2 | An application-level freeze: a local proxy stops forwarding bytes and closes nothing | 9 |
| 3 | Real packet loss: iptables drops the flow both ways in a Linux container, nothing is closed | 6 |
| Stage 2 | Round 3’s packet loss, with a client-side recovery | 11 |
The client detects a silent loss itself: a WebSocket ping every 0.5 s, and the link is lost after 2 s without a frame or a pong, which fired 1.5 to 2.0 s after the loss. The container ran on Colima started with --network-address --network-preferred-route, because its default network let a Mac process terminate Google’s TCP connection and keep acknowledging for a dead client.
What a resume keeps, and the 1011 right after an idle drop
A pending call survives a resume that works. In all 14 runs with a pending call and a successful resume, the resumed session accepted the FunctionResponse for the call id issued before the drop. No re-issued call, no toolCallCancellation, no error, one booking per run, and the call stayed answerable for as long as the session could be resumed: up to 154 s. The model only knows what the last tool response told it. In 3 of 3 runs where the client resumed but never sent the result, the model still said the booking was in progress, 2.9 to 3.1 s after the commit; in run 1, “I’m booking the 3 pm slot for you now.”
The immediate resume after an idle drop is refused. When the client aborted its connection with the model idle, a resume sent at once was closed with 1011 Internal error encountered. in 12 of 12 runs; one about 1.5 s after the drop worked in 12 of 12. With the model speaking at the drop, the immediate attempt worked in 3 of 3. That early 1011 is not a dead handle.
After a silent loss, the session stays locked
A silent loss is what a phone sees when it drops off Wi-Fi: packets stop, and no close frame, FIN or RST reaches the server. Round 3 dropped the live connection’s packets with iptables, while new connections passed.
results/B1.jsonl, B2.jsonl, BR1.jsonl.In those five runs, 450 of 450 resume attempts, made at detection and then every 10 s, were refused with the same 1011 for as long as I tried: 15 minutes. None hung. With the model idle, the server sent nothing for 479 s: no WebSocket ping, no TCP keepalive. With unacknowledged data in flight, its TCP gave up retransmitting after about 2 minutes, and the session stayed locked anyway. Round 2’s freeze showed the same: 40 of 40 attempts refused up to 133 s, nothing in a 12.8-minute probe.
What unlocks the session is a close. When the path came back and the client closed the old connection, the next resume worked 0.5 to 0.6 s later (6 of 6, rounds 2 and 3). The path coming back was not enough (N=1): with the old connection open again, attempts were refused until the client closed it.
The common real case is the one where no close can arrive: a phone that moves from Wi-Fi to cellular gets a new address and never reaches the server on the old connection again. In every silent-loss run, the booking committed 1 to 3 s after the loss. Where the old path never came back, the booking was made and the user was never told. The repository has a minimal reproduction for a laptop, and the lockout is filed as issue 60 on Google’s Live API examples repository.
A client-side recovery: four steps, the user told 6 to 7 s after the loss
recovery.py is a reference pattern for that case.
- Detect the loss with the ping rule above.
- Resume briefly: one attempt at detection, then one per second, none after 4 s, with a close queued on the old socket.
- Fall back to a new session without a handle and send one
send_client_content: the last turns of the client’s own transcript, then a “System note” with one status line per side effect, from the backend, not the model. The old call id is never answered. - Deduplicate: a
book_slotissued again is matched by business key and answered from the existing job.
In 9 of 9 recovered runs, the answer to “Did you book it?” matched the backend, with one booking each. The resume window never helped: 36 of 36 attempts were refused. With the status note, the model’s first sentence already said the booking was made, for example “Of course, your 3:00 p.m. slot for tomorrow has been successfully booked.”, and it never re-issued the call (0 of 7). From the loss to the new session’s first audio took 6.1 to 7.1 s, of which the resume window, useless on a dead path, took 3.55 to 3.6 s. The queued close never left the client. A plain resume loop without the fallback got 118 refusals out of 118 in 60 s and never told the user (N=2).
Without the note, only the dedupe prevented a double booking. In 2 of 2 runs that restored only the transcript, the model first re-issued book_slot, with a new id. The client matched the business key and answered from the committed job, so there was still one booking, but the model then said “I’m booking the 3:00 p.m. slot tomorrow for you.”, 3.6 to 3.8 s after the commit.
What a client must do
These follow from what I measured; they are not guarantees from the API.
- Keep a ledger of tool calls by call id across connections. After a resume, send the result for the old id, even if the side effect finished during the gap.
- Retry a refused resume after an idle drop: the 1011 cleared in about 1.5 s. Keep the handle you have; only one arrives per connection.
- Detect the loss yourself with pings; a healthy idle connection is silent for seconds.
- Get a close to the server when you can. If the path comes back, close the old connection before resuming.
- Put a short cut-off on resuming and fall back to a new session.
- In the new session, restore each side effect’s status from the backend and deduplicate by business key.
- Do not count on transparent mode or
goAway: the SDK refusestransparenton the Gemini Developer API, and nogoAwayarrived in any session.
What this test does not establish
- One endpoint: the Gemini Developer API through AI Studio, not Vertex AI (now Gemini Enterprise Agent Platform), whose documentation describes a transparent resumption mode.
- Synthetic speech, one macOS voice, and a fake in-process backend.
- Small N: 1 to 4 runs per scenario, one model, one network path, two days. It shows what can happen, not how often.
- No real device. The loss was iptables in a container behind three packet-level NATs (Docker’s, Apple’s, the home router’s), not a phone changing networks; the server saw a dead peer either way.
- One wording of the note, and the commit always landed inside the resume window.
- Not tested: a path that comes back after the server’s own close, a second loss after a fallback, a new session opened in parallel with the resume attempts.
Sources, code and raw data
The harness, recovery.py, the raw JSONL timelines, FINDINGS.md with every table and verbatim transcript, and the minimal reproduction are in frontier-on-cloud/gemini-live-resume-test at commit 687b1af; make_figure.py and make_clip.py in the same repository rebuild the figure and the clip from the JSONL. Google report: gemini-live-api-examples issue 60. Documentation, read on 2 October 2026: Session management with Live API, the Live API reference, Live API capabilities, Tool use with Live API, the Gemini 3.8 Live model page, and on Google Cloud, Start and manage live sessions and Best practices with Gemini Live API. The figure is at commit a2bc678; the clip and the two sessions recorded for it at 85a9bf8.
An AI coding agent wrote the harness under my direction; I am responsible for the design and for the claims in this article.