LLM systems · Real-time tool calls

When the connection drops mid-booking: session resumption on Gemini Live, measured

Part three of the Gemini Live field notes. A voice assistant loses its connection while a booking is pending. A resume that works keeps the call; after a silent loss, every resume was refused for 15 minutes, and only a new session got the truth to the user.

On this page

After “stop”, the connection itself

In part one, a user said “stop” while a gemini-3.8-live assistant was booking a slot, and no toolCallCancellation arrived. In part two, a client-side guard held the irreversible step until the application knew what the user had said. Both assumed the WebSocket stayed up, which it does not on a phone leaving Wi-Fi range. The Live API offers session resumption for this, but its documentation does not say what happens to a tool call pending at the disconnect, how soon a resume can succeed, or which errors it returns.

Four ways to lose a connection, 43 sessions

The assistant is the one from parts one and two: gemini-3.8-live on the AI Studio endpoint, google-genai 2.25.0, session resumption on from the start. A synthesized voice says “Book me the 3pm slot tomorrow, please.”, the model calls book_slot, and a fake backend commits 4 s later (7 s when the model is speaking). The connection is lost before the result reaches the model; the client resumes if it can, sends the result and asks “Did you book it?”, and the answer is compared with the backend. The sessions ran on 2 and 3 October 2026.

RoundHow the connection is lostSessions
1The client aborts its own TCP connection, so the server sees it close17
2An application-level freeze: a local proxy stops forwarding bytes and closes nothing9
3Real packet loss: iptables drops the flow both ways in a Linux container, nothing is closed6
Stage 2Round 3’s packet loss, with a client-side recovery11

The client detects a silent loss itself: a WebSocket ping every 0.5 s, and the link is lost after 2 s without a frame or a pong, which fired 1.5 to 2.0 s after the loss. The container ran on Colima started with --network-address --network-preferred-route, because its default network let a Mac process terminate Google’s TCP connection and keep acknowledging for a dead client.

What a resume keeps, and the 1011 right after an idle drop

A pending call survives a resume that works. In all 14 runs with a pending call and a successful resume, the resumed session accepted the FunctionResponse for the call id issued before the drop. No re-issued call, no toolCallCancellation, no error, one booking per run, and the call stayed answerable for as long as the session could be resumed: up to 154 s. The model only knows what the last tool response told it. In 3 of 3 runs where the client resumed but never sent the result, the model still said the booking was in progress, 2.9 to 3.1 s after the commit; in run 1, “I’m booking the 3 pm slot for you now.”

The immediate resume after an idle drop is refused. When the client aborted its connection with the model idle, a resume sent at once was closed with 1011 Internal error encountered. in 12 of 12 runs; one about 1.5 s after the drop worked in 12 of 12. With the model speaking at the drop, the immediate attempt worked in 3 of 3. That early 1011 is not a dead handle.

After a silent loss, the session stays locked

A silent loss is what a phone sees when it drops off Wi-Fi: packets stop, and no close frame, FIN or RST reaches the server. Round 3 dropped the live connection’s packets with iptables, while new connections passed.

Top: five runs with real packet loss on a log time axis up to 15 minutes. Every resume attempt, one every 10 s, is refused with 1011; the booking commits 1 to 3 s after the loss; the server's last TCP retransmission comes after about 2 minutes (model speaking) or 10 minutes (model idle); no resume is accepted. Bottom: four recovery runs. The loss is detected after 1.5 to 2 s, four resumes are refused in the window, a new session opens, and the model speaks 6.1 to 7.1 s after the loss, saying the booking is made in 4 of 4 runs.
Top: the five round-3 runs where the old path never came back (B1, model idle; B2, speaking). Bottom: the four BR1 recovery runs. Data: results/B1.jsonl, B2.jsonl, BR1.jsonl.

In those five runs, 450 of 450 resume attempts, made at detection and then every 10 s, were refused with the same 1011 for as long as I tried: 15 minutes. None hung. With the model idle, the server sent nothing for 479 s: no WebSocket ping, no TCP keepalive. With unacknowledged data in flight, its TCP gave up retransmitting after about 2 minutes, and the session stayed locked anyway. Round 2’s freeze showed the same: 40 of 40 attempts refused up to 133 s, nothing in a 12.8-minute probe.

What unlocks the session is a close. When the path came back and the client closed the old connection, the next resume worked 0.5 to 0.6 s later (6 of 6, rounds 2 and 3). The path coming back was not enough (N=1): with the old connection open again, attempts were refused until the client closed it.

The common real case is the one where no close can arrive: a phone that moves from Wi-Fi to cellular gets a new address and never reaches the server on the old connection again. In every silent-loss run, the booking committed 1 to 3 s after the loss. Where the old path never came back, the booking was made and the user was never told. The repository has a minimal reproduction for a laptop, and the lockout is filed as issue 60 on Google’s Live API examples repository.

A client-side recovery: four steps, the user told 6 to 7 s after the loss

recovery.py is a reference pattern for that case.

In 9 of 9 recovered runs, the answer to “Did you book it?” matched the backend, with one booking each. The resume window never helped: 36 of 36 attempts were refused. With the status note, the model’s first sentence already said the booking was made, for example “Of course, your 3:00 p.m. slot for tomorrow has been successfully booked.”, and it never re-issued the call (0 of 7). From the loss to the new session’s first audio took 6.1 to 7.1 s, of which the resume window, useless on a dead path, took 3.55 to 3.6 s. The queued close never left the client. A plain resume loop without the fallback got 118 refusals out of 118 in 60 s and never told the user (N=2).

Two sessions recorded on 3 October for this clip, not counted in the tables, with the user’s lines spoken by an open-source voice (Kokoro-82M, af_heart) and sent as real input. Without recovery: real packet loss, the booking commits, every resume is refused with 1011; the clip shows a 60 s excerpt with time compressed, and the 15-minute figure comes from the round-3 runs above. With recovery: four refusals, a new session with the status note, and the model’s own voice; its reply stops playing when the user asks again. Audio: the user speech as sent and the model’s output.

Without the note, only the dedupe prevented a double booking. In 2 of 2 runs that restored only the transcript, the model first re-issued book_slot, with a new id. The client matched the business key and answered from the committed job, so there was still one booking, but the model then said “I’m booking the 3:00 p.m. slot tomorrow for you.”, 3.6 to 3.8 s after the commit.

What a client must do

These follow from what I measured; they are not guarantees from the API.

What this test does not establish

Sources, code and raw data

The harness, recovery.py, the raw JSONL timelines, FINDINGS.md with every table and verbatim transcript, and the minimal reproduction are in frontier-on-cloud/gemini-live-resume-test at commit 687b1af; make_figure.py and make_clip.py in the same repository rebuild the figure and the clip from the JSONL. Google report: gemini-live-api-examples issue 60. Documentation, read on 2 October 2026: Session management with Live API, the Live API reference, Live API capabilities, Tool use with Live API, the Gemini 3.8 Live model page, and on Google Cloud, Start and manage live sessions and Best practices with Gemini Live API. The figure is at commit a2bc678; the clip and the two sessions recorded for it at 85a9bf8.

An AI coding agent wrote the harness under my direction; I am responsible for the design and for the claims in this article.

About the Frontier communities →

Keep exploring

Working on a similar problem? Let’s discuss it.

Get new articles via RSS