Where part one left off: the client is on its own
In part one, a scripted user asked a gemini-3.8-live assistant to book a slot through a tool, then said “stop” while the booking ran. No toolCallCancellation ever arrived, with text or with speech. The model’s statements about the booking were not reliable: with text it said the booking was already made before it was, and in BLOCKING mode with speech it said nothing was booked after a tool response reporting the booking. BLOCKING mode also booked twice. The fix has to live in the application, which holds the call id and runs the job.
Four behaviors around a two-phase service
The guard, commit_guard.py, sits in the application’s receive loop and sees every server message. It needs a service it can stop: prepare() does the slow, reversible part (here a wait standing in for an availability check and a hold), then either commit() performs the irreversible side effect or cancel() releases the job. The guard decides which of the two runs, and when.
Hold on speech onset. In part one, the earliest sign of a spoken stop was voiceActivity ACTIVITY_START, about 150 ms in. After prepare(), the guard waits a 1.5 s grace window and commits only if no user utterance is open. If the user speaks while a job is pending, the commit waits for that utterance’s transcript, and a keyword check (stop, cancel, don’t, wait and a few others) decides. The utterance that triggered the call is checked too, because a quick stop can reach the server merged with the request. With no transcript within 3 s, the default is to cancel. This answers the missing cancellation.
Deduplicate by business key. The key is the tool name plus the arguments reduced to what they mean: for book_slot, a day and a 24-hour time, tomorrow 15:00. A matching call within 30 s gets the existing job’s result and is never executed. An idempotency key built from the raw arguments would have missed: in one guarded BLOCKING run the first call said tomorrow at 3pm and the re-issue, with a new call id, tomorrow 3pm. Two of the seven re-issues in this test changed the wording. A real application would key on the resolved date-time or slot id. This answers the double booking.
Authoritative status note. After a cancel, or a commit that happened despite a stop, the guard sends the model a short user turn with send_client_content and turn_complete=True: “System note: the booking for tomorrow 3pm was cancelled before it was committed. Nothing is scheduled.” The FunctionResponse still carries the real status. This answers the unreliable statements.
Abandoned BLOCKING calls. In part one, after interrupted arrived during a pending BLOCKING call, the model acted as if the call no longer existed: it re-issued it, or ignored the response that reported the booking. The guard marks such a call abandoned, holds its commit and does not answer the dead id. Deduplication answers a re-issue, on its live id, with the existing job’s status. This answers both BLOCKING failures.
Unwanted bookings: 12 of 12 without the guard, 0 of 12 with it
The harness reuses part one’s setup: gemini-3.8-live on the AI Studio endpoint, google-genai 2.25.0, the same prompt and synthesized speech clips streamed in real time; the stop clip lasts 2.2 s. With the guard off, the service prepares and commits at once. On 1 October 2026, each scenario ran three times off and three times on. A: default mode, stop 1.0 s after the call, 4 s prepare. C: the same in BLOCKING mode. G2: a follow-up question makes the model speak while the call is pending, the stop comes 1 s into that speech, 7 s prepare. F: stop 0.3 s after the request ends, before any call, 4 s prepare. As in part one, no toolCallCancellation arrived in any of the 36 sessions.
results/summary.md.| Scenario | Unwanted commit, off → on | Double booking, off → on | Statement consistent, on |
|---|---|---|---|
| A, default mode | 3/3 → 0/3 | 0/3 → 0/3 | 3/3 |
| C, BLOCKING | 3/3 → 0/3 | 2/3 → 0/3 | 3/3 |
| G2, model speaking | 3/3 → 0/3 | 0/3 → 0/3 | 3/3 |
| F, stop before the call | 3/3 → 0/3 | 0/3 → 0/3 | 3/3 |
“Statement consistent” compares the transcript after the stop with the fake service. It flatters the guard-off runs, where “already made” was true, even when C booked twice. With the guard on, first and last claims were right in all 12 runs.
The hold prevented all 12 unwanted commits, in three ways. In A and C, the stop transcript arrived before the grace window ran out, so the window alone would have sufficed. In G2, it arrived 0.67 s after the window had ended: only the hold on the open utterance kept the commit back. In F, the check on the triggering utterance cancelled the job during prepare().
The decision comes about 3.5 s after the user starts saying stop: 3.46 to 3.51 s in A, C and G2, up to 3.83 s in F. That is the 2.2 s utterance plus the transcript, which arrived in one chunk with ACTIVITY_END, 1.26 to 1.31 s after the audio ended; no interim transcription arrived. The hold covers that gap from the onset on.
The hold protects the booking, the note protects the sentence
I ran two ablations the same morning, three runs each, with the guard’s runtime code unchanged.
Without the status note, the hold still prevented every commit. In A and G2, both in the default mode, the FunctionResponse saying cancelled went out 0.32 to 0.46 s before the model’s first words in A, and within 61 ms of them in G2. The first reply was false in all six runs. In A run 3, the model said “The booking was already made and cannot be canceled.”, then “I have stopped the process, and the slot was not booked.” In G2 run 2: “It’s too late; the booking was already made.”, then “Actually, the booking was cancelled and was never made.” Each correction began 2.4 to 4.8 s after the false reply. In C, the model re-issued the call in 3 of 3 runs, deduplication answered each new id with cancelled, and all three first replies were right: a response on a live BLOCKING call carried the status.
Without the hold and without abandon handling, which would otherwise hold a BLOCKING commit by itself, the commit in C went through in 3 of 3 runs. The guard’s note followed the stop transcript: “System note: booking BK-1001 for tomorrow 3pm is confirmed. The user asked to stop after it was committed. Tell them it is confirmed and offer to cancel it.” All three replies matched the committed state and offered to cancel; in run 1 the model said “The booking was already confirmed. Should I go ahead and cancel it for you?” The 12 guard-off sessions offered to cancel 0 times.
The two halves are separate. The hold protects the side effect; the note protects what the user is told. A tool response alone did neither: it cannot undo a commit, and in the default mode the model’s first reply contradicted it.
A status note can cut a reply in half
The note is a full user turn, and the Live guide says that turn_complete=true “unconditionally interrupts generation”. The model is often already answering the stop when the guard decides: in 8 of the 12 full-guard runs, interrupted arrived 12 to 78 ms after the note. In C run 3 the reply was split across two turns, “I have not” and “booked that slot for you.” Joined, the transcript is right. But none of the first turn’s audio had arrived, and a client that drops queued audio on interrupted, as Live clients do to stop talking over the user, would play only “booked that slot for you.”, the opposite of what happened. Two options, neither tested: send the note only while the model is idle, or carry the status in a FunctionResponse with a scheduling mode.
What this test does not establish
- The service is a fake:
prepare()is a sleep andcancel()always succeeds. A real backend needs its own reservation, expiry and idempotency. The grace window adds 1.5 s to every commit. - Speech was one synthetic voice, without noise or echo. With real microphones,
ACTIVITY_STARTcan fire on noise or on the model’s own audio; the guard would hold, then cancel without a transcript. - The intent check is a keyword list: “wait, make it 4pm” would cancel the booking.
- One model, one prompt and the AI Studio endpoint; no Vertex AI, no session resumption.
- N=3 per cell on one morning shows what can happen, not how often.
- Consistency is scored on transcripts, with rules tuned on 51 statements from this prompt. What a listener hears was not verified.
Most of this guard replaces signals the API does not send. A server-side cancellation for pending calls, or a documented signal that a BLOCKING call was dropped, would make most of it unnecessary.
Sources, code and raw data
The guard, the harness, the raw JSONL timelines and results/summary.md are in frontier-on-cloud/gemini-live-commit-guard at commit 20330de. ./run_guard.sh reruns the 24 before/after sessions with your own API key. Part one’s harness and data are in frontier-on-cloud/gemini-live-stop-test. Documentation: the Live API reference, the Live API tool use guide, the Live guide and the Gemini 3.8 Live model page.
An AI coding agent wrote the guard and the harness under my direction; I am responsible for the design and for the claims in this article.